NORTHLINE
← All recordsRECORD / 039
PaperNote

[Note] EvoRoute: Experience-Driven Self-Routing LLM Agent Systems

performance (i.e., task success and accuracy)

PaperNote9 min read

Agent System Trilemma

  • performance (i.e.i.e., task success and accuracy)
  • efficiency (i.e.i.e., time or steps required to complete tasks)
  • cost (i.e.i.e.., computational and monetary resources consumed)

EvoRoute

A self-evolving routing paradigm that dismantles the trilemma through fine-grained model selection. Before executing each step, it dynamically selects the most judicious LLM by:

  • retrieval: performing a multifaceted retrieval to identify historically analogous sub-task executions from an evolving knowledge base;
  • filtration: distilling a Pareto-optimal set of candidate models, i.e.i.e., those that are not dominated across the axes of cost, efficiency, and performance;
  • selection: leveraging a lightweight decision model to make the final selection based on this rich, context-aware statistical evidence.

Notations

M=⟨I,L,ϕ,S,T,A,Ψ,μ,Q⟩,\mathcal{M} = \langle \mathcal{I}, \mathcal{L}, \phi, \mathcal{S}, \mathcal{T}, \mathcal{A}, \Psi, \mu, \mathcal{Q} \rangle,
  • M\mathcal{M}: the designed complex agentic AI system
  • I={1,2,...,N}\mathcal{I} = \{1,2,...,N\}: the set of agent roles (e.g.e.g., web-browser, coder)
  • L\mathcal{L}: the pool of available LLM backbones
  • ϕ:I→L\phi:\mathcal{I}\rightarrow\mathcal{L}
  • S\mathcal{S} : system state, typically implemented as a shared memory or scratchpad
  • T\mathcal{T}: a set of external tools, such as code interpreters or web search APIs
  • A\mathcal{A}: the full action space, including both natural language actions and tool invocations, formally A=Alang∪{use_tool(T,args)∣T∈T}\mathcal{A}=\mathcal{A}_{lang}\cup\{\text{use\_tool}(T,args)\mid T\in\mathcal{T}\}
  • Ψ{st+1∣st,at}\Psi\{s_{t+1}\mid s_t,a_t\} governs the transition dynamics of the system
  • μ(t)∈I\mu(t)\in\mathcal{I} selects the active agent at each time step tt

Objective Formulation

ρ∗=arg⁡max⁡ρ(Eτ∼ρ[P(τ)],−Eτ∼ρ[C(τ)],−Eτ∼ρ[D(τ)])\rho^* = \arg\max_\rho \left( \mathbb{E}_{\tau \sim \rho}[\mathbb{P}(\tau)], -\mathbb{E}_{\tau \sim \rho}[\mathbb{C}(\tau)], -\mathbb{E}_{\tau \sim \rho}[\mathbb{D}(\tau)] \right)
  • ρ\rho: the dynamic routing policy that selects an LLM lt∈Ll_t\in\mathcal{L} for the active agent at each step tt
  • P(τ)\mathbb{P}(\tau): task performance
  • C(τ)\mathbb{C}(\tau): the cumulative monetary and computational expenditure
  • D(τ)\mathbb{D}(\tau): the total wall-clock execution time
  • τ=(s0,a0,s1,a1,...,sT)\tau=(s_0,a_0,s_1,a_1,...,s_T): the full execution trajectory of the system

Methodology

Step-level experience base

The backbone of EvoRoute is an evolving knowledge base K\mathcal{K} built from prior executions. After a task finishes, the full trajectory is split into step-level records:

Rt=⟨it,lt,qt,et,Tt,ct,dt,σt,P(τ)⟩,\mathcal{R_t} = \langle i_t, l_t, q_t, e_t, T_t, c_t, d_t, \sigma_t, \mathbb{P}(\tau) \rangle,

Each record stores:

  • iti_t: active agent role,
  • ltl_t: LLM used at this step,
  • qtq_t: sub-task instruction,
  • ete_t: embedding of the instruction,
  • TtT_t: tools used,
  • ctc_t: cost,
  • dtd_t: wall-clock duration,
  • σt\sigma_t: whether the step executed successfully,
  • P(τ)\mathbb{P}(\tau): final task-level success signal.

After each run:

K←K∪{Rt}t=0T−1\mathcal{K}\leftarrow\mathcal{K}\cup\{\mathcal{R}_t\}^{T-1}_{t=0}

Multi-Faceted Retrieval

When a new step arrives, EvoRoute retrieves relevant historical records from K\mathcal{K}. Instead of relying on one notion of similarity, it uses three.

  1. Agent Role Match

    Kagent={Rt∈K∣it=it′}\mathcal{K}_{\text{agent}}=\{\mathcal{R}_t\in\mathcal{K}\mid i_t=i_{t'}\}
  2. Semantic Similarity Retrieval

    Ksem={Rt∈K∣sim(Embed(qt′),et)≥θsim}\mathcal{K}_{\text{sem}} = \{\mathcal{R}_t \in \mathcal{K} \mid \text{sim}(\text{Embed}(q_{t'}), e_t) \geq \theta_{\text{sim}}\}
    • Embed(⋅)\text{Embed}(\cdot) is implemented via MiniLM
    • θsim=0.85\theta_{\text{sim}}=0.85
  3. Tool Congruence Retrieval

    Ktool={Rt∈K∣Tt∩PredictTools(qt′)≠∅}.\mathcal{K}_{\text{tool}} = \{\mathcal{R}_t \in \mathcal{K} \mid T_t \cap \text{PredictTools}(q_{t'}) \neq \emptyset\}.

    PredictTools(⋅)\text{PredictTools}(\cdot) uses a two-stage predictor:

    • Keyword heuristic
      • Using a predefined dictionary to map explicit trigger keywords(e.g., "search" for web_search\text{web\_search}; "run", "plot" for code_interpreter\text{code\_interpreter})
    • Cheap LLM fallback
      • if heuristics fail, use Qwen3-14B in zero-shot mode

The final candidate set:

Kcand=Kagent∪Ksem∪Ktool\mathcal{K}_{\text{cand}}=\mathcal{K}_{\text{agent}}\cup\mathcal{K}_{\text{sem}}\cup\mathcal{K}_{\text{tool}}

Pareto-Optimal Filtration and Selection

From the retrieved records, EvoRoute extracts candidate models:

Lcand={lt∣Rt∈Kcand}\mathcal{L}_{\text{cand}}=\{l_t\mid\mathcal{R}_t\in\mathcal{K}_{\text{cand}}\}

For each candidate model l∈Llandl\in\mathcal{L}_{\text{land}}, it estimates:

  • average performance P^(l)\hat{P}(l)
  • average cost C^(l)\hat{C}(l)
  • average delay D^(l)\hat{D}(l)

A model is dominated if another model exists that is superior or equal on all three axes and strictly superior on at least one.

Retaining only the non-dominated models, we form the Pareto-optimal set, Lpareto\mathcal{L}_{\text{pareto}}

Thompson-sampling-based model selection

If EvoRoute always picked the current best average, it would become too greedy and stop learning.

It assumes each metric follows a Normal distribution and models the uncertainty over its mean and variance using a Normal-Inverse-Gamma conjugate prior.

First, compute the sample statistics for each metric m∈{P,C,D}m \in \{\mathbb{P},\mathbb{C},\mathbb{D}\}: the count nln_l, the sample mean x‾m,l\overline{x}_{m,l} and the sample variance sm,l2s^2_{m,l}. These statistics are used to parameterize the NIG posteriors, NIG(μm,l,vm,l,αm,l,βm,l\mu_{m,l},v_{m,l},\alpha_{m,l},\beta_{m,l}), where μm,l=x‾m,l\mu_{m,l}=\overline{x}_{m,l}, vm,l=nlv_{m,l}=n_l, αm,l=nl/2\alpha_{m,l}=n_l/2, and βm,l=(nl−1)sm,l2/2\beta_{m,l}=(n_l-1)s^2_{m,l}/2

At decision time, it samples a stochastic utility:

U′(l)=wp⋅x~P,l−wc⋅x~C,l−wd⋅x~D,lU'(l) = w_p \cdot \tilde{x}_{P,l} - w_c \cdot \tilde{x}_{C,l} - w_d \cdot \tilde{x}_{D,l}

and selects:

l∗=arg⁡max⁡l∈Lpareto(U′(l)),l^* = \arg\max_{l \in \mathcal{L}_{\text{pareto}}} (U'(l)),

where (wp,wc,wd)(w_p,w_c,w_d) reflect the desired trilemma trade-off (wp=1.0w_p=1.0, wc=0.1w_c=0.1, wd=0.05w_d=0.05)

Crucially, this selection is not the end of the process. Once the agent powered by l∗l^∗ completes its action, the observed outcome is logged back into the knowledge base K\mathcal{K}. This closes the feedback loop, ensuring that every decision and its outcome contribute to the system’s ever-improving wisdom, thereby realizing the self-evolving nature of EvoRoute.

End of record / 039
← All records
READ NEXTWhat is Sepeculative Decoding