Instruments · Measurements
Measurements
Readings from the three instruments: prediction registry, tech-node status, action entries.
Demo data for layout preview only — replace the JSON files under src/data/ (predictions, tech-nodes, actions).
Forecasting
Prediction Registry
| Prediction | Prob. | Deadline | Status |
|---|---|---|---|
| Demo: an open-weight model reaches 85% on SWE-bench Verified | 65% | 2026-12-31 | Open |
| Demo: a major benchmark is shown to have large-scale contamination | 50% | 2026-09-30 | Open |
| Demo: a long-horizon agent works over 8 hours on a real repository | 40% | 2026-12-31 | Open |
| Demo: agent frameworks converge to 2–3 dominant paradigms | 70% | 2025-12-31 | Resolved · Hit |
Capability Sampling (inputs to forecasts)
6 entries
| Model | Org | Params | MMLU | MATH | Code |
|---|---|---|---|---|---|
| Qwen3-235B | Alibaba | 235B / 激活 22B | 87.8 | 85.0 | 70.7 |
| DeepSeek-V3.1 | DeepSeek | 671B / 激活 37B | 88.5 | 89.3 | 74.8 |
| Kimi-K2 | Moonshot AI | 1T / 激活 32B | 89.5 | 87.5 | 76.1 |
| GLM-4.6 | Zhipu AI | 355B / 激活 32B | 86.2 | 84.1 | 72.5 |
| Llama-4-Maverick | Meta | 400B / 激活 17B | 85.5 | 80.2 | 68.9 |
| Mistral-Medium-3 | Mistral | 未公开 | 84.1 | 78.6 | 66.3 |
Tech Tree
Node Status
| Node | Dependencies | Status |
|---|---|---|
| Efficient sparse attention | Prior sparse-attention work, hardware support | Unlocked |
| Long-horizon agent memory | Efficient sparse attention | Bottleneck (key node) |
| Trustworthy autonomous research | Agent memory, trusted eval protocols | Locked |
Atlas
Action Entries
| Goal | Actor | Configuration |
|---|---|---|
| Reduce benchmark contamination | Leaderboard maintainers | Public protocols + canary items |
| Improve evaluation reproducibility | Researchers | Open-source scripts + fixed seeds |
| Make forecasts accountable | Media & analysts | Prediction registry + public calibration |