Releases: kosuke1701/AnimeCV
Release list
Yet Another Character Face Annotations on Danbooru2020
Description
I applied the pre-trained face detection model in AnimeCV to the SFW 512px downscaled subset of Danbooru2020 dataset.
Applied model is FaceDetector_EfficientDet(coef=2).
It contains 6,412,982 face annotations for 3,227,706 imges.
How to use
Information of extracted face bounding boxes are stored in Sqlite3 database file faces.sql.
Example code
import os
import sqlite3
import matplotlib.pyplot as plt
import numpy as np
from PIL import Image
# Target image
filename = "/data/danbooru2020/512px/512px/0841/2543841.jpg"
name = f"danbooru/{os.path.basename(filename)}"
# Retrieve face position
conn = sqlite3.connect("faces.sql")
c = conn.cursor()
c.execute("SELECT * FROM faces WHERE name=?", (name,))
lst_faces = c.fetchall()
# Crop and show face image
img = Image.open(filename)
for i, (_id, name, xmin, ymin, xmax, ymax) in enumerate(lst_faces):
ax = plt.subplot(1, len(lst_faces), i+1)
crop_img = img.crop((xmin, ymin, xmax, ymax))
ax.imshow(np.asarray(crop_img))
plt.show()Database schema
table name: faces
| column | type | description |
|---|---|---|
| id | INTEGER | Unique id for each row. |
| name | STRING | Name of the image. danbooru/<basename of original image file>. (e.g.) danbooru/4039070.jpg |
| xmin | INTEGER | Left coordinate |
| ymin | INTEGER | Top coordinate |
| xmax | INTEGER | Right coordinate |
| ymax | INTEGER | Bottom coordinate |
Pre-trained anime character identification model
Near Human-Level Character Identification model (Updated on 2021-02-07)
I trained models on ZACI-20 dataset derived from Danbooru 2020, which is much larger than my private dataset used to train the previous model (0111_best_randaug).
Performance on the test-set of ZACI-20 is as follows. The best model (0206_resnet152) achieves near human-level performance with an error rate (FPR) only 1.5 times greater than that of humans.
Note that this performance is evaluated on images of novel characters which are not contained in the training dataset.
| model name | FPR (%) | FNR (%) | EER (%) | note |
|---|---|---|---|---|
| Human | 1.59 | 13.9 | N/A | by kosuke1701 |
| ResNet-152 | 2.40 | 13.9 | 8.89 | w/ RandAug, Contrastive loss. 0206_resnet152 |
| SE-ResNet-152 | 2.43 | 13.9 | 8.15 | w/ RandAug, Contrastive loss. 0206_seresnet152 |
| ResNet-152 | 2.54 | 13.9 | 8.33 | w/ RandAug, Contrastive + Classification (Cross Entropy) loss. 0301_cls_resnet152 |
| ResNet-18 | 2.96 | 13.9 | 8.65 | w/ RandAug, Contrastive + Classification (Cross Entropy) loss. 0217_cls_resnet18 (Updated on 2021-02-18) |
| ResNet-18 | 5.08 | 13.9 | 9.59 | w/ RandAug, Contrastive loss. 0206_resnet18 |
0111_best_randaug
I trained ResNet-18 model with Contrastive loss and data augmentation by RandAugment. Output embedding is 500-dimensional vector with unit L2 norm. Similarity metric used in training is L2 Euclidian distance.
Performance of zero-shot character identification on my private dataset is as follows.
| method | False positive rate (FPR) | False negative rate (FNR) | Equal error rate (EER) |
|---|---|---|---|
| Human (me) | 1.22% | 22.0% | N/A |
| 0111_best_randaug | 11.3% | 22.0% | 14.3% |
Other notes:
- Demo code is available:
- Training code is available at https://github.com/kosuke1701/optuna-metric-learning