Skip to content
Demonstrator · items marked Example are invented · what exists today
E2ER

External pipelines · RISE project catalogue

MLGym (Meta)

A gym-style framework and benchmark (MLGym-Bench, arXiv:2502.14499) for advancing AI research agents on 13 diverse ML research tasks (CV, NLP, RL, game theory). Like [`aviary`](aviary.md), MLGym is **evaluation infrastructure** for RISE-style systems rather than a RISE pipeline itself.

Indexed in RISE · dormantConformance with the standard plannedProject site

What it does

First gym environment specifically for *ML research tasks* — not general QA or knowledge work — covering idea generation, data processing, method implementation, training, and result analysis as a unified RL training surface. Distinct from Aviary in that it targets the *training* of research agents via RL, not only their evaluation.

Focus
end-to-end
Inputs
task-specification
Outputs
agent-trajectories, evaluation-metrics, training-data
Architecture
tool-use, artifact-versioning, iterative-loop
Maintained by
Meta AI / FAIR
Started
2025

Description

Data model
Discipline
Computer science
Method family
not specified
Design
not specified
Research stage
HypothesesResearch designData analysisCode generation
Contributors
Meta AI / FAIR
Usage
not used in published research yet
Source
RISE project catalogue · projects/landscape · @4c17bae
Record
pipeline:mlgym · JSON

Solid tags are declared by the source or mapped from its terms; dashed tags are inferred by a published rule. Hover a tag for its provenance.

Bring it into the standard

A pipeline built outside E2ER can meet the standard by describing its steps as a template, attaching the floor of checks and publishing evaluation records. Its authors keep ownership and credit.

Other pipelines in Computer science