OmnilingualGAIA2: Evaluating AI Agents Across Ten Languages
A machine-translated expansion of four GAIA2 capabilities, with seven agents evaluated in ten languages beyond English.
Read more →
A machine-translated expansion of four GAIA2 capabilities, with seven agents evaluated in ten languages beyond English.
Read more →I'm a PhD student at Meta Superintelligence Labs in Paris, supervised by Thomas Scialom (Meta) and Djamé Seddah (Inria Paris, ALMAnaCH team). My research focuses on building evaluation frameworks for LLM agents — designing realistic, dynamic environments that test how well AI systems can operate autonomously in the real world.
My work includes ARE, a scalable framework for agent evaluation environments; GAIA2, an ICLR 2026 Oral benchmark for dynamic and asynchronous tasks; and OmnilingualGAIA2, which measures how agentic competence transfers beyond English.
I studied Computer Science (ML specialization) at Georgia Tech (2023–2025) and Engineering in Computer Science at Université de Technologie de Compiègne (2019–2024).
A machine-translated expansion of four GAIA2 capabilities across ten target languages, with a calibrated verifier and an analysis of the pooled performance gap beyond English.
Evaluating AI agents in dynamic, event-driven scenarios that mirror real-world complexity.
A framework for building diverse, scalable agent evaluation environments.
A machine-translated expansion of four GAIA2 capabilities, with seven agents evaluated in ten languages beyond English.
First-class terminal support for Gaia2 — evaluate agents like OpenClaw and Hermes-Agent in real execution environments.
New evaluation results on Gaia2 — Claude 4 Sonnet Extended Thinking, DeepSeek, Qwen, and more. Analysis of model coverage and community reception.
Co-authored with Clémentine Fourrier. An overview of the GAIA2 benchmark and the ARE framework, and how the community can use them.
If you'd like to discuss my research, potential collaborations, or anything else — feel free to reach out!