GIFT-EVAL / 2026-09-14yoft-o
The orchestration model behind yoft
The lowest normalized MASE among forecasting methods that use neither dataset-specific fine-tuning nor LLMs, based on the published GIFT-Eval results.
Among methods without fine-tuning or LLMs#3 overall, without these conditions†
- Point forecast error
- 0.650MASE ↓
- Probabilistic forecast error
- 0.444CRPS ↓
- Entries in the full benchmark
- 129entries
‡ Our ranking by normalized MASE uses the public results retrieved on 2026-09-14. “No fine-tuning” means no dataset-specific training of model weights. “No LLM” means no LLM in the forecasting pipeline, including retrieval selection. EXAONE-Forecast-Agent and STRIDE w/ Synapse, the two higher-scoring entries, are excluded because their primary sources document LLM use in forecasting. Under these conditions, yoft-o (listed as TIMEHEDGE) ranks first. This is not an official leaderboard category. The assistant for explaining results and conversation is a separate feature.
† Based on 129 GIFT-Eval entries retrieved on 2026-09-14, ranked by normalized MASE. This differs from the leaderboard’s default ordering. Lower MASE and CRPS are better.
These results evaluate the model used for regularly sampled forecasts. Performance on public datasets does not guarantee accuracy on your data.