Radar Brief Week 25, 2026 · 2026-08-29 — 2026-09-05

September 3: BAAI releases 3 embodied datasets in quick succession
human judgment becomes a bottleneck for robot collaboration

This week scanned 86 HF orgs · 50 GitHub orgs · 71 blogs · 125 X accounts

0
Valuable Datasets
0
Related Papers
0
Blog Posts
0
Active Repos
One-line Summary

Agent-related GitHub repositories continued expanding as of 2026-09-05 [P0], BAAI released multi-robot and embodied planning data on 2026-09-03 [P0], and NVIDIA plus LAION turned code Agent trajectories into the main battleground [P1]. The strongest data demand signal in this scan: code Agent trajectories.

Key Findings

This week's 5 high commercial value findings

P0 Agent-related GitHub repositories continued expanding as of 2026-09-05 [P0]

As of 2026-09-05, NousResearch/hermes-agent has 241,719 stars, up 22,688 from 2026-07-23. openai/codex has 121,647 stars, up 20,903 from 2026-07-23. anthropics/skills has 174,335 stars, up 10,853 from 2026-07-23. anthropics/claude-code has 144,124 stars, up 5,374 from 2026-07-23. anthropics/claude-plugins-official has 35,926 stars, up 3,416 from 2026-07-23.

Business significance → The Agent toolchain has shifted from showcasing model capabilities to reusable workflows and a skills marketplace. This will directly raise demand for human judgment data on tool-use trajectories, success/failure determination, replay-based fixes, memory summarization, and more. Knowlyr can prioritize the high-value scenario of human reviewing Agent trajectories, so judgment contributions can generate revenue around task completion quality, executability, and safety.
P0 BAAI released multi-robot and embodied planning data on 2026-09-03 [P0]

BAAI/Discoverse-L was released on 2026-09-03, with 129 downloads. BAAI/MobileVLA-CoT was released on 2026-09-03, with 56 downloads. BAAI/Orchestra-Bench was released on 2026-09-03, with 186 downloads, containing 12,000 samples and aimed at three-robot collaborative planning. BAAI/ToolPrivBench was released on 2026-08-31, with 53 downloads, focusing on how agents choose between "standard tools" and "high-privilege tools".

Business significance → Chinese research institutions are pushing embodied AI data from single-robot actions toward multi-view perception, collaborative planning, permission selection, and chain-of-thought reasoning. What is most lacking here is not more synthetic trajectories, but expert judgment on which subtasks are more reasonable, which tools should be authorized, and which collaboration plans are safer. For data service companies, this is a high-ARPU human-machine collaborative review opportunity.
P1 NVIDIA and LAION turn code Agent trajectories into the main battleground [P1]

NVIDIA/Open-SWE-Traces currently has 21,253 downloads, up from 8,274 on 2026-07-23, a 156.9% increase. This repository was released on 2026-04-16, and the v1.2 update added agent trajectories generated by Qwen3.8-27B. On the LAION side, multiple terminal_bench- and swebench-related datasets were added this scan, including laion/terminal_bench_2_a3_rl_DCAgent_exp_rpt_unitsyn_python_v3_10_8B_20260829_192924, laion/terminal_bench_2_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260f49e33da, laion/swebench_verified_random_100_folders_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_c2a08420, and others.

Business significance → Code Agents no longer need only static benchmarks; they need trajectory-level data — the entire process from environment interaction and tool calls to failure recovery and final submission. This direction naturally depends on human judgment to distinguish "apparently successful" from "truly shippable," and also on human labeling-based decisions about error types, fix quality, and safety risks. Whoever can deliver high-quality trajectory reviews can enter the Agent training supply chain.
P1 Evaluation and preference learning signals clustered from 2026-09-01 to 2026-09-03 [P1]

allenai/BenchMIRT was released on 2026-09-01, with 327 downloads. allenai/BenchMIRT-item-statistics has 352 downloads. allenai/BenchMIRT-model-statistics has 342 downloads. allenai/BenchMIRT-eval-data has 39 downloads. Meanwhile, the paper "Subspace Inference Enables Efficient Active Reward Learning from Preferences" was published on 2026-09-03, "Patterning in Practice: Debiasing Reward Models with Susceptibilities" was published on 2026-09-01, and "Small Language Models as Judges for Rubric-Based Reinforcement Learning" came into community view on 2026-08-30.

Business significance → Reward models and preference learning are moving from batch scoring to actively selecting the most informative samples. This means the most valuable asset in the future will not just be large-scale preference data, but high-quality human judgment that can identify disputed, boundary, and biased samples. Knowlyr should focus on judge-style data, especially in scenarios that require expert consistency, conflict arbitration, and rubric-level review.
P2 Multilingual speech and cultural understanding are entering a tradable data-supply phase [P2]

google/WaxalNLP was released on 2026-01-19, with 25,857 downloads and 274 likes, and is a large-scale multilingual speech corpus for African languages. The paper "TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding" was published on 2026-09-01, targeting Persian dialogue tasks for more than 120 million users. The paper "MemeBridge: A Dataset for Benchmarking and Mitigating the Bidirectional Cultural Gap in Meme Interpretation" was published on 2026-08-31, focusing on cross-cultural meme understanding.

Business significance → Multilingual and cultural-context data are upgrading from language coverage to cultural understanding and expression quality. Such data depends heavily on native speakers and cultural experts, especially for speech naturalness, dialogue politeness, semantic ambiguity, and interpretation of cultural metaphors. For a model that insists on "letting people earn income through contributed judgments," this is the long-term track that best reflects the irreplaceability of humans.

Demand Signals

Infer training data demands from model releases

Data Type Intensity Trend Related Signals
Code Agent trajectories
very strong ↑ New
NVIDIA/Open-SWE-Traces downloads reached 21,253, while LAION added terminal_bench and swebench trajectory data in this scan
Multi-robot/embodied planning data
very strong ↑ New
BAAI/Orchestra-Bench · MobileVLA-CoT · NVIDIA/PhysicalAI-Robotics-Open-H-Embodiment are advancing in parallel
Agent tool use and permission selection
very strong ↑ New
LAION agent_tool categories grew from 5 to 20 versus 2026-07-23, and BAAI/ToolPrivBench · Snowflake/HybridDeepResearch emerged
Preference learning / reward model data
strong ↑ New
BenchMIRT · Subspace Inference · Patterning in Practice · Small Language Models as Judges appeared in concentration
Multilingual speech and dialogue data
strong ↑ New
google/WaxalNLP 25,857 downloads, TalkFa · Ready to Speak appeared at the same time
Synthetic data for security and identity verification
strong ↑ New
IDSPACE · ToolPrivBench · Google/DeepMind cyber defense-related signals intensified
Long-document and knowledge-work evaluation
medium ↑ New
Microsoft/XL-DocBench · RHELM · Asta citation data remain active
Cultural / multimodal understanding data
medium ↑ New
MemeBridge · SnapBench · video_benchmarks point to cross-cultural and mobile multimodal tasks
Time-series forecasting data
medium ↑ New
google/timesfm-3.0-pytorch 123,025 downloads, with weather and industrial forecasting blogs linked
robot manipulation trajectories ↓ Dropped Present in previous issue, absent this issue
real Agent trajectories and long-horizon evaluation ↓ Dropped Present in previous issue, absent this issue
open-ended generative rubrics ↓ Dropped Present in previous issue, absent this issue
scientific analysis evaluation data ↓ Dropped Present in previous issue, absent this issue
first-person video memory / 3D reasoning ↓ Dropped Present in previous issue, absent this issue
long-video audio/video temporal localization ↓ Dropped Present in previous issue, absent this issue
Agent executable skill library ↓ Dropped Present in previous issue, absent this issue
synthetic environments and world model data ↓ Dropped Present in previous issue, absent this issue
document robustness and confidence calibration ↓ Dropped Present in previous issue, absent this issue
expert gold sets and contamination-resistant evaluation ↓ Dropped Present in previous issue, absent this issue

Download Movers

Datasets with the largest download changes this week

Dataset Downloads Weekly Growth
lerobot/community_dataset_v3 32,028 +526.3%
allenai/asta-bench-submissions 215 +186.7%
nvidia/Open-SWE-Traces 21,253 +156.9%
allenai/asta-summary-citation-counts 1,095 -22.8%

Want to discuss this issue?

Kai
Kai Founder & CEO
苏文
苏文 AI Documentation & Release Engineer
陆明哲
陆明哲 AI Product Manager

Auto-generated by AI Dataset Radar · Updated weekly

AI Dataset Radar →