issai/Wild1024
Viewer β’ Updated β’ 1.02k β’ 26
Predicts font family (75), language (11), color (64, EGA index), and
style (4) from a text-image crop. Backbone: google/siglip2-base-patch16-naflex
(native aspect ratio, max_num_patches=256); pooled feature (d=768) from the
SigLIP2 attention-pooling head feeds four independent heads
(Linear(768β512) β LayerNorm β GELU β Dropout(0.1) β Linear(512βn)).
| font | language | color | style |
|---|---|---|---|
| 96.34 | 96.19 | 95.92 | 96.98 |
Wild1024 out-of-distribution: language 82.42, font-category 89.55.
from huggingface_hub import snapshot_download
import sys, torch, json
from transformers import AutoProcessor
from PIL import Image
path = snapshot_download("issai/FontID")
sys.path.insert(0, path)
from modeling_fontid import FontIDModel
model = FontIDModel.from_pretrained(path).eval()
proc = AutoProcessor.from_pretrained("google/siglip2-base-patch16-naflex")
maps = json.load(open(f"{path}/class_mappings.json"))
img = Image.open("crop.png").convert("RGB")
o = proc(images=img, max_num_patches=256, return_tensors="pt")
with torch.no_grad():
out = model(o["pixel_values"], o["pixel_attention_mask"], o["spatial_shapes"])
for head in ["font", "lang", "color", "style"]:
idx = int(out[head].argmax(-1))
print(head, maps[head][str(idx)])
Paper is coming soon.
Base model
google/siglip2-base-patch16-naflex