feat(moa): per-reference-model max_tokens override

MoA reference_max_tokens is preset-level — one cap for all reference
models. When mixing a verbose model with a terse one, a single cap is
either too tight for the terse model or too loose for the verbose one.

Now each reference slot can optionally carry its own max_tokens:

  reference_models:
    - provider: openrouter
      model: deepseek/deepseek-v4-pro
      max_tokens: ***        # per-slot cap, overrides preset-level
    - provider: openai-codex
      model: gpt-5.5
      # no max_tokens → falls back to preset-level reference_max_tokens

_clean_slot (moa_config.py) preserves an optional max_tokens field on
the slot dict, coerced via _coerce_int_or_none. _run_reference
(moa_loop.py) reads slot-level max_tokens first, falling back to the
preset-level cap passed by the caller. Slots without the field are
unaffected — backward compatible.

Type hints on slot-handling functions updated from dict[str, str] to
dict[str, Any] to reflect the now-heterogeneous slot shape.
This commit is contained in:
Rain 2026-07-07 18:18:23 +02:00 • committed by Teknium
parent ead9d7b256
commit bc7212cf93
5 changed files with 178 additions and 5 deletions

View file

@ -105,6 +105,14 @@ def _clean_slot(slot: Any) -> dict[str, Any] | None:
effort = _clean_reasoning_effort(slot.get("reasoning_effort"))
if effort:
clean["reasoning_effort"] = effort
# Optional per-slot max_tokens: overrides the preset-level
# reference_max_tokens for this specific reference model. None (the
# default) = no cap, so existing slots are unaffected. Allows tuning
# each advisor's output length independently — useful when one model
# is verbose and another is terse.
slot_mt = _coerce_int_or_none(slot.get("max_tokens"))
if slot_mt is not None:
clean["max_tokens"] = slot_mt
return clean