Flame-Chasers is an open-source research organization working on person retrieval and re-identification at the intersection of computer vision, vision-language learning, and multimodal interaction.
Our central question is simple: how can an intelligent system find the right person from complex visual collections using the information people naturally provide? We study this question through language descriptions, cross-modal representations, synthetic supervision, multilingual queries, challenging viewpoints, dialogue, and multimodal feedback.
Rather than treating person retrieval as a fixed image-matching problem, we aim to build systems that understand fine-grained identity cues, learn with less manual annotation, adapt to diverse scenarios, and collaborate with users throughout the search process.
We develop fine-grained representations that connect visual appearance with natural-language descriptions. Our work explores semantic alignment, relation reasoning, sensitivity-aware learning, and vision-language pretraining to distinguish subtle identity-related details in large galleries.
High-quality person descriptions are expensive to collect and may raise privacy concerns. We investigate learning with limited or non-parallel supervision, generated descriptions, fully synthetic data, and large-scale pseudo-labeled resources. We also extend retrieval beyond conventional settings through multilingual descriptions and aerial-ground viewpoints.
A single query rarely captures everything a user knows. We are moving person retrieval from one-shot matching toward multi-turn interaction, where dialogue, follow-up questions, corrective feedback, and retrieved reference images progressively refine the search. This direction brings together cross-modal alignment, conversational memory, and adaptive visual understanding.
Our repositories are designed to turn research ideas into usable resources. Alongside implementations, we release training and evaluation recipes, model checkpoints, annotations, datasets, and empirical benchmarks whenever possible, making it easier to reproduce results and build upon them.
Semantic alignment → Vision-language foundations → Data-efficient learning → Conversational retrieval → Multimodal interaction
Across this journey, the goal remains the same: to make person retrieval more accurate, flexible, accessible, and aligned with the way humans communicate.
- Official implementations of our research in text-based, chat-based, and interactive person retrieval
- Reproducible pipelines for supervised, unsupervised, synthetic-data, multilingual, and cross-view settings
- Public datasets, annotations, pretrained models, checkpoints, and evaluation tools
- Ongoing explorations of foundation models and multimodal agents for human-centered visual search
Browse our open-source projects to find code, datasets, models, and detailed usage instructions. Questions, discussions, and research collaborations are welcome through the issue tracker of the most relevant project.
We encourage responsible research and applications that respect privacy, consent, fairness, and applicable laws.