Add any Hugging Face GGUF model, not just the curated catalog - #38
Add any Hugging Face GGUF model, not just the curated catalog#38sachin-detrax wants to merge 2 commits into
Conversation
|
Strong PR, thank you. The input parser is the best part of it: repo URLs, Two blocking items, both small:
Nice to have, your call whether in this PR or as follow-ups:
Two things we will split into separate issues rather than grow this PR: memory-requirement estimates read from the GGUF header instead of the file-size heuristic (you left a note about this yourself), and hardware-aware quant selection instead of the constant Q4 preference. This PR does not need to carry them. With 1 and 2 in, this is good to merge from my side. |
The local model list was a fixed array of nine models compiled into
models-catalog.ts. Anything else on Hugging Face was unreachable without
running llama-server by hand or waiting for a maintainer to ship a new
release.
Models are now open-ended. A user-added entry is a full LocalModelDef
with a `custom-` prefixed id, persisted in localModels.customModels
(config v33 -> v34, older files inherit []). loadConfig registers them
in a catalog registry, so getLocalModelDef / isKnownLocalModelId resolve
curated and custom entries through one lookup and every existing
consumer -- daemon start, installer, TUI rows, CLI -- picked them up
without changes.
Entry points:
TUI "+ Add a model from Hugging Face..." pinned under Local text
models. Enter opens a prompt taking either a reference (resolve
and add) or free text (search, then pick by digit).
CLI models add <ref> / models search <query>.
/models add and /models search from the prompt.
References accepted: repo URLs, /resolve/ and /blob/ file URLs,
hf://owner/repo[@rev]/file.gguf, a pasted `hf download ...` command
(one- or two-argument, trailing flags dropped), and bare owner/name.
Naming only a repo picks a 4-bit quant and any mmproj projector, which
makes vision repos come out vision-capable. HF_TOKEN is honoured for
gated repos, including on the weights download.
hf:// is parsed by hand rather than with `new URL`: that puts the owner
in the host slot and lowercases it, and HF owners are case-sensitive, so
hf://Qwen/... would silently 404.
Two rendering defects surfaced while testing this, both pre-existing and
both able to shred the panel on any daemon failure:
- daemonError / errorLine are view-model fields rendered as a single
<Text> in a one-row slot, but a failed llama-server start stored a
4KB, ~25-line log tail. Ink overlaps lines when a frame outgrows its
budget rather than clipping: 48 rows into a 24-row budget. Flattened
on write and on read, with wrap="truncate-end" on the one-row slots
-- flattening alone still wrapped a long single line.
- llm-panel computed its list budget without subtracting the modals
and banners it renders above the list, unlike local-models-panel:
42 rows into 24, and 66 with both defects together.
Flattening the message hid the real cause behind "did not become
healthy", so extractLoadFailure now lifts the diagnostic line out of the
log tail into the message head. Full detail still reaches the runtime
feed and the L logs tab.
Routing between "add this model" and "set this base URL" goes through
looksLikeHuggingFaceReference, stricter than the parser: a bare
owner/name is ambiguous with 192.168.1.5/api and stays behind the
explicit `add` verb. An inline regex here had already drifted once --
it knew about huggingface.co but not hf://, so /models hf://... was
silently persisted as http://hf://... instead of adding the model.
Tests: HF reference parsing and quant/projector selection, add-row
placement and prompt key handling (including panel hotkeys not stealing
characters mid-URL and digits only picking when results show), rendered
frames, and frame-height regressions asserting height <= budget.
Closes AtomicBot-ai#27
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…iles The two blocking review items on AtomicBot-ai#38: - any shard of a multi-part GGUF (first included) is now rejected with a clear error at pick time, instead of downloading one part and failing only at llama-server launch; repos shipping only sharded weights get a dedicated message pointing at single-file quants - speculative-decoding companion files (MTP/NextN) are excluded from the smallest-file fallback and rejected on an explicit pick, so the default quant selection can never land on a non-runnable GGUF Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2beefa0 to
9385f5e
Compare
|
Rebased onto main per the author's request (config bump lands as v36, after stopOnExit and progressIndicator took 34 and 35) and picked up the two blocking review items: sharded picks are now rejected outright with a clear error (first part included, plus a dedicated message for sharded-only repos), and MTP/NextN companion files are excluded from the default-quant fallback. All 62 tests from the original branch pass unchanged, 4 new ones cover the guards. Ready for review. |
Closes #27.
The local model list was a fixed array of nine models compiled into
models-catalog.ts. Anything else on Hugging Face was unreachable without runningllama-serverby hand or waiting for a maintainer to ship a new release. Models are now open-ended.A user-added entry is a full
LocalModelDefwith acustom-prefixed id, persisted inlocalModels.customModels(config v33 → v34; older files inherit[]).loadConfigregisters them in a catalog registry, sogetLocalModelDef/isKnownLocalModelIdresolve curated and custom entries through one lookup — daemon start, installer, TUI rows and CLI all picked them up without changes.Entry points
+ Add a model from Hugging Face..., pinned under Local text models. Enter opens a prompt taking either a reference (resolve and add) or free text (search, then pick by digit).models add <ref>andmodels search <query>./models addand/models search.References accepted
/resolve/and/blob/file URLshf://owner/repo[@rev]/file.ggufhf download ...command (one- or two-argument, trailing flags dropped)owner/nameNaming only a repo picks a 4-bit quant and any mmproj projector, so vision repos come out vision-capable.
HF_TOKENis honoured for gated repos, including on the weights download.hf://is parsed by hand rather than withnew URL: that puts the owner in the host slot and lowercases it, and HF owners are case-sensitive, sohf://Qwen/...would silently 404.Two pre-existing rendering defects fixed along the way
Both surfaced while testing this, and both could shred the panel on any daemon failure:
toStatusLine()truncates at every error-message write and at the render point, withwrap="truncate-end"on the InkText.LlmPanelframe budget did not account for overlay modal / banner rows, so opening a modal during an error overran the frame.estimateOverlayRows()now covers them.Regression tests pin the frame height for both.
Also
extractLoadFailurepulls the diagnostic reason out of llama-server logs so a failed load reports why rather than a generic message.Tests
35 files changed, +2104/−52, including new suites for the HF reference parser, panel rendering, frame height, and the reducer.
The 6 failures are identical by name on
mainand on this branch (ChatLog/SplashBanner/TuiAppsmoke,llm-panel selectors,persistEmbeddingHybridRecall) — pre-existing, untouched by this change. 38 tests added, none broken.🤖 Generated with Claude Code