hermes-agent/tests/plugins
emozilla 83f88909ef feat(desktop): first-class local Ollama provider with detection, model management, and capability-aware picking
Promote bare "ollama" from a custom-endpoint alias to a real provider and
build the desktop UX around it. A local server's connection kind is
reachability rather than a credential, so every credential-shaped gate
(provider registry, picker filters, settings surfaces) gets an explicit
path for it.

Backend:
- Provider overlay + registry entry (127.0.0.1:11434/v1 default, keyless
  with a local-only placeholder, base-url normalization for /api and /v1
  forms). Existing provider=custom configs are untouched.
- GET /api/local-servers/detect fingerprints well-known local ports plus a
  local configured base_url; response shape leaves room for a future
  installed/running/managed distinction.
- /api/ollama/* management endpoints: installed+running+recommended models,
  registry pull as a poll-able background job streaming native NDJSON
  progress, delete, and load (warm-up / keep_alive pinning). Pull and
  delete bust the picker's model-id cache.
- Model picker payload: per-model capabilities widened to tools/vision/
  context_length; local Ollama rows enriched from the server's native
  /api/show (authoritative for on-disk tags, where models.dev is sparse),
  backfilled off the request path by a background thread.
- The explicit-only picker filter keeps ollama rows: the row only exists
  when the server answered a probe, which is as explicit as a pasted key.
- Reasoning safety: /api/show thinking capability gates all reasoning
  fields (Ollama 400s reasoning_effort on non-thinking models), and
  OpenAI-only effort levels map to the nearest accepted level
  (xhigh->max, minimal->low).
- model.ollama_keep_alive config: sent per-request as extra_body.keep_alive.
- Latency discipline for a local server that may be down: a 300ms TCP
  pre-check with a short negative cache guards every native-API read; the
  status endpoint probes fresh so a just-started server is noticed
  immediately; localhost is rewritten to 127.0.0.1 (Windows resolves
  localhost to ::1 first and Ollama binds IPv4 loopback — each request
  otherwise pays a ~2s failed IPv6 connect, including chat inference).

Desktop:
- Providers -> Accounts: a "Local servers" card mirroring the OAuth card
  language ("Running · N models" / a start-the-server hint), expanding to
  model management: installed models with size/quant/VRAM, delete, warm-up,
  curated pull recommendations with a progress bar, free-form pull, and a
  KV-cache advisory when a loaded model runs well under its trained window.
  Polls while down so it flips to Running by itself.
- Model picker: "No tools" badge (explicit tools:false only — absence means
  unknown) with demotion, plus context window / parameter size / quant per
  row; refetches once after open so backfilled metadata appears in place.
- Onboarding: detected-server row with a model select, replacing blind
  first-model assignment for detected servers.
- Composer status stack: "Loading <model> into memory" row during cold
  starts, confirmed against /api/ps so ordinary slow generations stay quiet.

Chat inference stays on the OpenAI-compatible /v1 endpoint; the native
/api surface is used read-only for metadata plus explicit management
actions. Lifecycle management (starting or installing Ollama) is not
included.
2026-07-09 13:50:40 -04:00
..
browser fix(browser): self-review pass — dead-import, log levels, future-proofing 2026-05-17 04:04:15 -07:00
dashboard_auth feat(dashboard_auth): support confidential clients (client_secret) in self-hosted OIDC (#55344) 2026-06-30 13:32:51 +10:00
image_gen fix(image-gen): route local-input credential guard through one shared chokepoint + cover xai (#57698) 2026-07-03 18:47:53 +05:30
memory feat(mem0): add self-hosted mode to the setup wizard 2026-07-07 14:07:16 -07:00
model_providers feat(desktop): first-class local Ollama provider with detection, model management, and capability-aware picking 2026-07-09 13:50:40 -04:00
platforms/photon fix(photon): auto-reinstall stale sidecar deps before start 2026-07-05 17:38:32 -07:00
transcription feat(stt): add stt.providers.<name> command-provider registry 2026-05-25 01:41:19 -07:00
tts feat(tts): add register_tts_provider() plugin hook (closes #30398) 2026-05-24 18:04:54 -07:00
video_gen fix(xai): route video-gen local inputs through the shared read guard 2026-07-03 18:47:53 +05:30
web revert(web): remove keyless Parallel search fallback (#46350) 2026-06-14 16:47:57 -07:00
__init__.py fix: mem0 API v2 compat, prefetch context fencing, secret redaction (#5423) 2026-04-05 22:43:33 -07:00
test_achievements_plugin.py test: use subprocesses for each test file (#29016) 2026-05-21 16:40:04 +05:30
test_chronos_cron.py fix(cron): avoid provider package shadowing core cron 2026-06-23 23:39:22 -07:00
test_chronos_verify.py fix(cron): avoid provider package shadowing core cron 2026-06-23 23:39:22 -07:00
test_discord_runtime_failure.py fix(discord): recover from runtime gateway task exits (#44383) 2026-06-11 15:39:01 -04:00
test_disk_cleanup_plugin.py fix: protect cron output root from cleanup 2026-07-01 15:42:04 +05:30
test_google_meet_audio.py chore: prune unused imports and duplicate import redefinitions 2026-05-28 22:26:25 -07:00
test_google_meet_node.py chore: prune unused imports and duplicate import redefinitions 2026-05-28 22:26:25 -07:00
test_google_meet_plugin.py test(google_meet): assert ladder-based dependency install instead of bespoke pip argv 2026-07-07 04:09:35 -07:00
test_google_meet_realtime.py chore: prune unused imports and duplicate import redefinitions 2026-05-28 22:26:25 -07:00
test_hindsight_health_grace_timeout.py feat(hindsight): configurable embedded daemon health grace timeout (#50341) 2026-06-21 12:20:53 -07:00
test_hindsight_root_guard.py fix(hindsight): skip local_embedded daemon when running as root 2026-06-21 11:47:02 -07:00
test_kanban_attachments.py feat(kanban): file attachments on tasks (#35395) 2026-05-30 07:41:04 -07:00
test_kanban_dashboard_plugin.py fix(security): sanitize kanban markdown html 2026-06-21 13:10:17 -07:00
test_kanban_worker_runs.py feat(kanban): add POST /runs/{run_id}/terminate endpoint 2026-05-29 00:21:54 -07:00
test_langfuse_plugin.py test(langfuse): pin exact surviving key in turn-isolation test 2026-06-18 13:00:01 +05:30
test_nemo_relay_plugin.py fix(nemo-relay): preserve downstream errors in adaptive execution (#42691) 2026-06-09 02:31:10 -07:00
test_plugin_dashboard_auth_contract.py fix(dashboard): sanction plugin WS/upload auth via SDK helpers (gated mode) 2026-06-03 16:59:36 -07:00
test_raft_check_fn_silent.py fix(plugins): silence raft check_fn log spam for users without raft CLI 2026-06-19 17:12:58 -07:00
test_retaindb_plugin.py chore: prune unused imports and duplicate import redefinitions 2026-05-28 22:26:25 -07:00
test_security_guidance_plugin.py chore: prune unused imports and duplicate import redefinitions 2026-05-28 22:26:25 -07:00
test_teams_pipeline_plugin.py fix(teams-pipeline): reject dot-only recording display_name 2026-07-01 02:03:48 -07:00