Romain Froger

Romain Froger

PhD Student · Meta Superintelligence Labs, Paris
Meta Inria
Building agents that work in the real world
New Blog Post

OmnilingualGAIA2: Evaluating AI Agents Across Ten Languages

A machine-translated expansion of four GAIA2 capabilities, with seven agents evaluated in ten languages beyond English.

Read more →

About

I'm a PhD student at Meta Superintelligence Labs in Paris, supervised by Thomas Scialom (Meta) and Djamé Seddah (Inria Paris, ALMAnaCH team). My research focuses on building evaluation frameworks for LLM agents — designing realistic, dynamic environments that test how well AI systems can operate autonomously in the real world.

My work includes ARE, a scalable framework for agent evaluation environments; GAIA2, an ICLR 2026 Oral benchmark for dynamic and asynchronous tasks; and OmnilingualGAIA2, which measures how agentic competence transfers beyond English.

I studied Computer Science (ML specialization) at Georgia Tech (2023–2025) and Engineering in Computer Science at Université de Technologie de Compiègne (2019–2024).

Publications

Preprint · 2026

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

Andrea Caciolai, Pere-Lluís Huguet Cabot, Chierh Cheng, Albert Ventayol-Boada, Gabriel Mejia Gonzalez, Christophe Ropers, Lucas Bandarkar, Sebastian Ruder, Darlene Sakakihara, Elliot Yun, Pierre Andrews, Grégoire Mialon, Romain Froger, Marta R. Costa-jussà

A machine-translated expansion of four GAIA2 capabilities across ten target languages, with a calibrated verifier and an analysis of the pooled performance gap beyond English.

ICLR 2026 · Oral

Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments

Romain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja, Ricardo Silveira Cabral, Virginie Do, Emilien Garreau, Jean-Baptiste Gaya, Hugo Laurençon, Maxime Lecanu, Kunal Malkan, Dheeraj Mekala, Pierre Ménard, Gerard Moreno-Torres Bertran, Ulyana Piterbarg, Mikhail Plekhanov, Mathieu Rita, Andrey Rusakov, Vladislav Vorotilov, Mengjue Wang, Ian Yu, Amine Benhalloum, Grégoire Mialon, Thomas Scialom

Evaluating AI agents in dynamic, event-driven scenarios that mirror real-world complexity.

Preprint · 2025

ARE: Scaling Up Agent Environments and Evaluations

Romain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja, Ricardo Silveira Cabral, Virginie Do, Emilien Garreau, Jean-Baptiste Gaya, Hugo Laurençon, Maxime Lecanu, Kunal Malkan, Dheeraj Mekala, Pierre Ménard, Gerard Moreno-Torres Bertran, Ulyana Piterbarg, Mikhail Plekhanov, Mathieu Rita, Andrey Rusakov, Vladislav Vorotilov, Mengjue Wang, Ian Yu, Amine Benhalloum, Grégoire Mialon, Thomas Scialom

A framework for building diverse, scalable agent evaluation environments.

Blog

2026 Blog Post

OmnilingualGAIA2: Evaluating AI Agents Across Ten Languages

A machine-translated expansion of four GAIA2 capabilities, with seven agents evaluated in ten languages beyond English.

2026 Blog Post

Gaia2-CLI: Evaluating Agentic Systems in Real Execution Environments

First-class terminal support for Gaia2 — evaluate agents like OpenClaw and Hermes-Agent in real execution environments.

2025 HuggingFace

Gaia2 Leaderboard Update: New Models and New Observations

New evaluation results on Gaia2 — Claude 4 Sonnet Extended Thinking, DeepSeek, Qwen, and more. Analysis of model coverage and community reception.

2025 HuggingFace

Gaia2 and ARE: Empowering the community to study agents

Co-authored with Clémentine Fourrier. An overview of the GAIA2 benchmark and the ARE framework, and how the community can use them.

Contact

If you'd like to discuss my research, potential collaborations, or anything else — feel free to reach out!