HN
Today

Show HN: Open-source model routing for coding agents at Astra-level performance

Weave Router 2.0, an open-source model router, intelligently orchestrates multiple LLMs for coding agents, achieving Astra-level performance at half the cost and double the speed. HNers are intrigued by its technical approach—using HMMs and smart caching—to overcome single-model limitations and optimize resource use in AI development. This story highlights a practical, cost-effective advancement in AI tooling.

69
Score
20
Comments
#8
Highest Rank
5h
on Front Page
First Seen
Oct 1, 6:00 PM
Last Seen
Oct 1, 10:00 PM
Rank Over Time
108101516

The Lowdown

Weave Router 2.0, an open-source model router, has achieved a significant milestone in AI agent performance. It intelligently switches between large language models (LLMs) for coding tasks, aiming to outperform any single model through an ensemble approach.

  • The router plugs into existing coding agents (e.g., Claude Code, Codex) and dynamically selects the best LLM for a given task, such as using Astra for debugging and Deepseek for simpler updates.
  • Benchmarked against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas, Weave Router 2.0 showed equivalent pass rates.
  • It achieved these results at approximately 50% of Astra's cost and 2.2x to 2.5x faster task completion.
  • The core improvements stem from three areas: a new architecture using a Hidden Markov Model (HMM) and classifier to manage the vast search space of routing decisions, a larger training dataset bootstrapped with frontier LLMs, and a smarter cache-eviction impact calculation to minimize unnecessary model switches and associated costs.
  • The project is open-source (GitHub link provided) and also offers a hosted version.

This breakthrough validates the hypothesis that an ensemble of models can surpass the capabilities of any single frontier model, offering a cost-effective and performant solution for coding agents.

The Gossip

Routing Ruminations and Refinements

Users inquired about the technical specifics of the router's operation, particularly how it handles model selection, provider variance (e.g., OpenRouter), and the definition and resilience of "model buckets." There was also curiosity about its ability to route to local/LAN-hosted models and comparisons to similar tools like Cursor's or Copilot's auto mode. The author clarified that for their hosted version, they carefully select providers and implement logic for deprioritizing degraded ones, while OpenRouter simplifies setup for self-hosting despite potential variance. They explained that model buckets are based on similar capabilities and that incorrect routing decisions can be "escalated" and corrected, though with a performance penalty.

Performance Ponderings and Price Points

A key discussion point revolved around how the router, trained on data labeled by frontier models, can *exceed* their performance, especially beyond just cost savings. The author responded by emphasizing that different frontier models excel at different tasks, and the router's strength lies in optimally combining them. There was also a question about predefined budget support, to which the author confirmed that this feature is indeed available.

Openness and Operational Questions

Users asked if the trained routing model itself is available as open weights, to which the author stated it is not. There was also interest in whether the router can route to locally or LAN-hosted open-weight models, and the author confirmed that this capability is supported.