BAAI Adds Three Embodied Datasets
human judgment Shifts Toward Robot Collaboration and Tool Safety
This week scanned 86 HF orgs · 50 GitHub orgs · 71 blogs · 125 X accounts
AllenAI released the BenchMIRT trio in one go on 2026-09-01 [P0], BAAI rolled out embodied and tool-safety data in a concentrated push from 2026-09-02 to 2026-09-03 [P0], and OpenAI brought code, cybersecurity, and "verifiability" to the forefront on 2026-09-03 [P0]. This week’s strongest data demand signal: robot collaboration / embodied planning data.
Key Findings
This week's 5 high commercial value findings
On 2026-09-01, allenai/BenchMIRT-item-statistics had 352 downloads and 1 like; allenai/BenchMIRT-model-statistics had 342 downloads and 0 likes; allenai/BenchMIRT had 327 downloads and 0 likes. The companion allenai/BenchMIRT-eval-data was released on 2026-08-25 and had 39 downloads. The official description explicitly states that this benchmark is used to measure LLM latent safety and general reasoning scores, and that the data includes prompt and model response statistics.
On 2026-09-02, BAAI/ToolPrivBench had 53 downloads. On 2026-09-03, BAAI/Orchestra-Bench had 186 downloads, and the description states that it includes 12,000 samples, collaborative planning for three robots, three-view inputs per sample, and three coordinated subtasks; BAAI/Discoverse-L had 129 downloads; BAAI/MobileVLA-CoT had 56 downloads.
On 2026-09-03, OpenAI released GPT-6 Astra and a Safety overview. The official statement said it reached a new level in computer use, coding, cybersecurity, and science, and classified its safety capability as Critical under the Preparedness Framework. On the same day, Daybreak for Frontline Defenders was also released, announcing a $1 billion investment to support frontline defenders. In the accompanying GitHub ecosystem, openai/codex has 121,712 stars, openai/evals has 19,387 stars, and openai/openai-agents-python has 29,208 stars.
On 2026-09-03, Subspace Inference Enables Efficient Active Reward Learning from Preferences discussed how to actively learn reward models from preferences. On 2026-09-01, Patterning in Practice: Debiasing Reward Models with Susceptibilities proposed using susceptibilities to debias reward bias. On 2026-09-01, StudentSim: Training LLM-based Student Simulators pointed out that real learner feedback is scarce and expensive. On 2026-09-03, Instruction Duplication as an Inference-Time Control Primitive showed that inference-time control can also serve as a structured intervention mechanism.
On 2026-08-30, EleutherAI/djinn-problems-v1.0 had 95 downloads, and the description explicitly states that it is a dual-verifier reward-hacking environment; the fixed-djinn v2 build was updated again on 2026-09-04. In the same period, multiple retrain bank releases went live: EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004 had 406 downloads, seed1006 had 383, seed1007 had 272, seed1005 had 174, and seed1008 had 531; EleutherAI/PARTIAL_LDS-retrain-bank-gpt2medium-16k-bs32 had 184 downloads.
Demand Signals
Infer training data demands from model releases
Want to discuss this issue?
Auto-generated by AI Dataset Radar · Updated weekly
AI Dataset Radar →