NeurIPS 2026 · Evaluations & Datasets Track

Re:Cognize

Open-Set Comic Character Re-Identification

Aaditya Baranwal, Madhav Kataria*, Yogesh S. Rawat, Shruti Vyas

Institute of Artificial Intelligence, University of Central Florida

*Work done as an intern at the University of Central Florida.

Recognising characters while the story is read. Left: a reader's question, such as where Takagi asks Mashiro, depends on who appears on which page. Centre: standard Re-ID matches each query against a gallery built before reading starts, so a new character finds no match. Right: Re:Cognize streams crops in reading order against four galleries: every character (P1), a few labelled examples of each (P2), none (P3), or a few that grow as it reads (P4). Re:Cast grows the gallery only where additions are right more often than the gallery on the queries they take over.
protocols on one query stream
4
protocols on one query stream
Re-ID backbones evaluated
5
Re-ID backbones evaluated
corpora: POPCharacters, Manga109, Re:Verse
3
corpora: POPCharacters, Manga109, Re:Verse
points of top-1 that correct labels would add
20+
points of top-1 that correct labels would add

01 · Abstract

Reading along, not matching against a cast list

A manga reader meets a character on one page and knows them on sight a hundred pages later, without ever being handed a cast list. Re-identifying comic characters demands the same, open-set and sequential: pages arrive as a stream in reading order, new faces appear before anyone names them, and the cast is assembled as the story is read. Re:Cognize evaluates recognition as the story is read, not against a cast handed over in advance: four protocols on one query stream, from closed-set retrieval to a cast the model must build and grow itself.

The surprise is where models fail. Recognising is close to solved: one reference image per character already ranks as well as a gallery built in advance. Knowing what to believe is not: a model that adds its own matches makes its cast worse, while the same growth with correct labels would gain over twenty points of top-1 accuracy. The bottleneck is acceptance, not vision, and one comparison decides it: an addition pays exactly when it is right more often than the cast already was on the queries it takes over. The comparison has nothing to fit, and measured on half of a new corpus it calls the other half correctly.

Re:Cast puts it to work with nothing fitted on data: a cast sheet of one running average per character, grown only where the page itself vouches for a crop. It recovers a third to two thirds of what perfect labels would, depending on whether the cast starts from random examples or from first appearances. Re:Cognize measures whether a model can read along; Re:Cast is a cast that does.

02 · Key findings

Where models fail

Recognising a character is close to solved. Deciding which of your own matches to believe is not.

01

Recognising is close to solved

One reference image per character already ranks as well as a gallery built in advance.

02

Knowing what to believe is not

A model that adds its own matches makes its cast worse, while the same growth with correct labels would gain over twenty points of top-1 accuracy.

03

One comparison decides it

An addition pays exactly when it is right more often than the cast already was on the queries it takes over. The comparison has nothing to fit, and measured on half of a new corpus it calls the other half correctly.

04

Re:Cast reads along

One running average per character, grown only where the page itself vouches for a crop, recovers a third to two thirds of what perfect labels would.

“Re:Cognize measures whether a model can read along; Re:Cast is a cast that does.”

03 · The framework

Four protocols, one stream

Every protocol answers the same stream of query crops in reading order; they differ only in the gallery and whether it may change.

ONE QUERY STREAM · BAKUMAN CH. 1 · READING ORDER →p. 8p. 17p. 25p. 29p. 43p. 44p. 47p. 50P1full gallery–P2one seed each–P3starts emptyjoins at ≥ 0.55cluster 1cluster 2cluster 3–P4grows by top-1–
readyTakagiMashiroAzuki
One stream, four galleries. MagiV2’s own decisions (its released encoder) on eight crops of Bakuman chapter 1, against the galleries of the paper’s protocols figure; the protocols run on whole chapters. The page-29 Mashiro crop sits closer to Azuki’s seed: P1 still gets it right, P2 does not, and P4 files it under Azuki, where two later Mashiro crops match it.
P1

Closed-set retrieval

A fifth of each character's crops form the gallery and the rest are queries, ranked by cosine similarity and scored by mAP and Rank-k. The closed-set ceiling.

P2

Seeded static gallery

k labelled crops per character, drawn at random (Seq-R) or as the character's first k appearances (Seq-T), as a reader meets them; every other crop is a query.

P3

Online clustering

Crops arrive unlabelled into an empty gallery and join the nearest cluster above a novelty threshold or open a new one. A diagnostic of emergence: one fixed threshold compares every representation.

P4

Seeded gallery that grows

P2's gallery, but in reading order each query is added under its top-1 match. Static, predicted and oracle policies separate what growth gives, scored by identity Rank-1.

04 · The commit condition

When a change to the gallery pays

The bottleneck is acceptance, not vision, and one comparison decides it.

A change to the gallery can alter the answer only for a query whose nearest entry it supplied: the queries it captures. Over the stream, the change in accuracy is exactly

Δ = c (peff − a+)

c

Capture rate

The share of queries whose nearest entry the change supplied.

peff

Precision where it acts

The share of captured queries whose capturing entry carries their own identity.

a+

The gallery it replaces

The accuracy the unchanged gallery would have had on those same queries.

So a change pays exactly when peff > a+: when its additions are right more often than the gallery already was on the queries they take over, not on average. A crop filed under the right character can still take over other characters’ queries, which is why neither term is the intuitive one.

05 · Re:Cast

A cast that reads along

A cast sheet of one running average per character, grown only where the page itself vouches for a crop, with nothing fitted on data.

ONE QUERY STREAM · BAKUMAN CH. 1 · READING ORDER →p. 8p. 17p. 25p. 29p. 43p. 44p. 47p. 50P4grows by its own top-1–Re:Castone entry each, page-checkedthe seed alonethe seed alonethe seed alone–
readyTakagiMashiroAzuki
Growth by top-1 against Re:Cast, on the same stream. Both start from one seed per character. P4 files every query under its top-1 match, so the page-29 mistake stays in the gallery and is matched again. Re:Cast keeps one averaged entry per character and adds a crop only when MagiV2’s page grouping ties it to a crop already committed, here the seeds on pages 17 and 50; everywhere else it holds.
Re:Cast on Re:Zero crops from Re:Verse, with Rom as identity A. Top: the three changes. A page names a character when one of its crops is already committed to it, and commitment then adds that crop's page-group sibling. Bottom: over six pages the cast sheet is updated only on the four that name Rom; elsewhere the rule abstains.
a

A cast sheet, not a bag of crops

Each character is one ℓ2-normalised average of the crops filed under it, so a new crop refines its character's entry instead of becoming a rival entry that takes over other characters' queries.

b

Commitment under the page constraint

A crop joins a character only when another crop in its page group is already committed to that character. The rule consults only crops already read, and abstains wherever the page offers no evidence.

c

Seed expansion

Before the stream starts, the crops on a seed's page that share its character join it: 1.43 per seed on average, 90.6 % of them correct.

Chronological seeding

Binding, and pricing the binder

When the seeds are each character’s first appearances, binding supplies the reference: page groups merged in reading order and named by their first seed, committed only where the commit condition predicts a gain.

Binding, and pricing the binder. Top, on Re:Verse's Re:Zero annotations: (a) a page's crops are grouped above one similarity threshold, (b) page groups merge in reading order into the best earlier match above a second, or start new ones, and (c) a group's first seed names it. Bottom: any frozen encoder can bind, with threshold τ; its crops are committed when the commit condition predicts a positive Δ, and τ is chosen by that prediction, never by the measured gain.

06 · Results

What the protocols measure, and where the headroom is

Five backbones, two corpora: one seed recovers the closed-set mAP, correct growth would add over twenty points, and the commit condition says which changes pay.

aWhat one seed recovers, by metric

bWhat correct growth adds, k = 1

cEncoder-side adaptation

dEvery change to the gallery, k = 5

What the protocols measure, and where the headroom is. (a) One seed per character recovers the closed-set mAP but only part of its Rank-1. (b) Growth by the model's own top-1 matches stays below a static gallery, while correct labels would add over twenty points (labels: oracle minus static). (a, b): memory-block configuration, one random seed per character, three training runs. (c) Encoder-side adaptation helps most on the backbones weakest in this domain. (d) Each point is one change to the gallery on one backbone and corpus: adding under the page constraint, adding by top-1 match, or restricting the candidate characters. The stronger the gallery already is, the less any change adds, and the commit condition predicts which points fall below zero.

08 · Citation

Cite Re:Cognize

BibTeX
@inproceedings{baranwal2026recognize,
  title     = {Re:Cognize: Open-Set Comic Character Re-Identification},
  author    = {Baranwal, Aaditya and Kataria, Madhav and Rawat, Yogesh S. and Vyas, Shruti},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Evaluations and Datasets Track},
  year      = {2026}
}