Skip to content

Releases: kosuke1701/AnimeCV

Yet Another Character Face Annotations on Danbooru2020

Choose a tag to compare

@kosuke1701 kosuke1701 released this 18 Jan 09:37

Description

I applied the pre-trained face detection model in AnimeCV to the SFW 512px downscaled subset of Danbooru2020 dataset.
Applied model is FaceDetector_EfficientDet(coef=2).

It contains 6,412,982 face annotations for 3,227,706 imges.

How to use

Information of extracted face bounding boxes are stored in Sqlite3 database file faces.sql.

Example code

import os
import sqlite3

import matplotlib.pyplot as plt
import numpy as np
from PIL import Image

# Target image
filename = "/data/danbooru2020/512px/512px/0841/2543841.jpg"
name = f"danbooru/{os.path.basename(filename)}"

# Retrieve face position
conn = sqlite3.connect("faces.sql")
c = conn.cursor()
c.execute("SELECT * FROM faces WHERE name=?", (name,))
lst_faces = c.fetchall()

# Crop and show face image
img = Image.open(filename)

for i, (_id, name, xmin, ymin, xmax, ymax) in enumerate(lst_faces):
    ax = plt.subplot(1, len(lst_faces), i+1)

    crop_img = img.crop((xmin, ymin, xmax, ymax))

    ax.imshow(np.asarray(crop_img))
plt.show()

Database schema

table name: faces

column type description
id INTEGER Unique id for each row.
name STRING Name of the image. danbooru/<basename of original image file>. (e.g.) danbooru/4039070.jpg
xmin INTEGER Left coordinate
ymin INTEGER Top coordinate
xmax INTEGER Right coordinate
ymax INTEGER Bottom coordinate

Pre-trained anime character identification model

Choose a tag to compare

@kosuke1701 kosuke1701 released this 11 Jan 09:40

Near Human-Level Character Identification model (Updated on 2021-02-07)

I trained models on ZACI-20 dataset derived from Danbooru 2020, which is much larger than my private dataset used to train the previous model (0111_best_randaug).

Performance on the test-set of ZACI-20 is as follows. The best model (0206_resnet152) achieves near human-level performance with an error rate (FPR) only 1.5 times greater than that of humans.

Note that this performance is evaluated on images of novel characters which are not contained in the training dataset.

model name FPR (%) FNR (%) EER (%) note
Human 1.59 13.9 N/A by kosuke1701
ResNet-152 2.40 13.9 8.89 w/ RandAug, Contrastive loss. 0206_resnet152
SE-ResNet-152 2.43 13.9 8.15 w/ RandAug, Contrastive loss. 0206_seresnet152
ResNet-152 2.54 13.9 8.33 w/ RandAug, Contrastive + Classification (Cross Entropy) loss. 0301_cls_resnet152
ResNet-18 2.96 13.9 8.65 w/ RandAug, Contrastive + Classification (Cross Entropy) loss. 0217_cls_resnet18 (Updated on 2021-02-18)
ResNet-18 5.08 13.9 9.59 w/ RandAug, Contrastive loss. 0206_resnet18

0111_best_randaug

I trained ResNet-18 model with Contrastive loss and data augmentation by RandAugment. Output embedding is 500-dimensional vector with unit L2 norm. Similarity metric used in training is L2 Euclidian distance.

Performance of zero-shot character identification on my private dataset is as follows.

method False positive rate (FPR) False negative rate (FNR) Equal error rate (EER)
Human (me) 1.22% 22.0% N/A
0111_best_randaug 11.3% 22.0% 14.3%

Other notes: