Recognising is close to solved
One reference image per character already ranks as well as a gallery built in advance.
Open-Set Comic Character Re-Identification
Aaditya Baranwal, Madhav Kataria*, Yogesh S. Rawat, Shruti Vyas
Institute of Artificial Intelligence, University of Central Florida
*Work done as an intern at the University of Central Florida.
01 · Abstract
A manga reader meets a character on one page and knows them on sight a hundred pages later, without ever being handed a cast list. Re-identifying comic characters demands the same, open-set and sequential: pages arrive as a stream in reading order, new faces appear before anyone names them, and the cast is assembled as the story is read. Re:Cognize evaluates recognition as the story is read, not against a cast handed over in advance: four protocols on one query stream, from closed-set retrieval to a cast the model must build and grow itself.
The surprise is where models fail. Recognising is close to solved: one reference image per character already ranks as well as a gallery built in advance. Knowing what to believe is not: a model that adds its own matches makes its cast worse, while the same growth with correct labels would gain over twenty points of top-1 accuracy. The bottleneck is acceptance, not vision, and one comparison decides it: an addition pays exactly when it is right more often than the cast already was on the queries it takes over. The comparison has nothing to fit, and measured on half of a new corpus it calls the other half correctly.
Re:Cast puts it to work with nothing fitted on data: a cast sheet of one running average per character, grown only where the page itself vouches for a crop. It recovers a third to two thirds of what perfect labels would, depending on whether the cast starts from random examples or from first appearances. Re:Cognize measures whether a model can read along; Re:Cast is a cast that does.
02 · Key findings
Recognising a character is close to solved. Deciding which of your own matches to believe is not.
One reference image per character already ranks as well as a gallery built in advance.
A model that adds its own matches makes its cast worse, while the same growth with correct labels would gain over twenty points of top-1 accuracy.
An addition pays exactly when it is right more often than the cast already was on the queries it takes over. The comparison has nothing to fit, and measured on half of a new corpus it calls the other half correctly.
One running average per character, grown only where the page itself vouches for a crop, recovers a third to two thirds of what perfect labels would.
“Re:Cognize measures whether a model can read along; Re:Cast is a cast that does.”
03 · The framework
Every protocol answers the same stream of query crops in reading order; they differ only in the gallery and whether it may change.
A fifth of each character's crops form the gallery and the rest are queries, ranked by cosine similarity and scored by mAP and Rank-k. The closed-set ceiling.
k labelled crops per character, drawn at random (Seq-R) or as the character's first k appearances (Seq-T), as a reader meets them; every other crop is a query.
Crops arrive unlabelled into an empty gallery and join the nearest cluster above a novelty threshold or open a new one. A diagnostic of emergence: one fixed threshold compares every representation.
P2's gallery, but in reading order each query is added under its top-1 match. Static, predicted and oracle policies separate what growth gives, scored by identity Rank-1.
04 · The commit condition
The bottleneck is acceptance, not vision, and one comparison decides it.
A change to the gallery can alter the answer only for a query whose nearest entry it supplied: the queries it captures. Over the stream, the change in accuracy is exactly
Δ = c (peff − a+)
c
Capture rate
The share of queries whose nearest entry the change supplied.
peff
Precision where it acts
The share of captured queries whose capturing entry carries their own identity.
a+
The gallery it replaces
The accuracy the unchanged gallery would have had on those same queries.
So a change pays exactly when peff > a+: when its additions are right more often than the gallery already was on the queries they take over, not on average. A crop filed under the right character can still take over other characters’ queries, which is why neither term is the intuitive one.
05 · Re:Cast
A cast sheet of one running average per character, grown only where the page itself vouches for a crop, with nothing fitted on data.
Each character is one ℓ2-normalised average of the crops filed under it, so a new crop refines its character's entry instead of becoming a rival entry that takes over other characters' queries.
A crop joins a character only when another crop in its page group is already committed to that character. The rule consults only crops already read, and abstains wherever the page offers no evidence.
Before the stream starts, the crops on a seed's page that share its character join it: 1.43 per seed on average, 90.6 % of them correct.
Chronological seeding
When the seeds are each character’s first appearances, binding supplies the reference: page groups merged in reading order and named by their first seed, committed only where the commit condition predicts a gain.
06 · Results
Five backbones, two corpora: one seed recovers the closed-set mAP, correct growth would add over twenty points, and the commit condition says which changes pay.
aWhat one seed recovers, by metric
bWhat correct growth adds, k = 1
cEncoder-side adaptation
dEvery change to the gallery, k = 5
07 · Code & data
The harness, the per-tuple results and the three corpora. With the results archive, every table and figure of the paper regenerates without a GPU.
The evaluation harness with all four protocols, Re:Cast, the memory-block baseline, and the analyses behind every number in the paper.
github.com/Per-tuple results behind every table and figure (23 MB). scripts/fetch_results.py downloads and verifies them; the tables then regenerate without a GPU.
recognize-results.tar.xzThe public PopCharacters subset of PopManga: 23 series, with 4,058 test crops of 70 characters over 8 held-out series.
huggingface.co/27 held-out volumes (784 characters, 29,315 crops), evaluated zero-shot.
www.manga109.orgThe Re:Zero annotations behind the Re:Cast and binding figures.
re-verse.vercel.app08 · Citation
@inproceedings{baranwal2026recognize,
title = {Re:Cognize: Open-Set Comic Character Re-Identification},
author = {Baranwal, Aaditya and Kataria, Madhav and Rawat, Yogesh S. and Vyas, Shruti},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Evaluations and Datasets Track},
year = {2026}
}