Changelog
All notable changes to ComfyUI Smart LM Loader are documented in this file.
Entries follow conventional commit prefixes:
2026-09-19
Version: 1.0.19
Fix
Restore Smart LM Loader frontend registration, keep Loader and Smart Detection metadata bound to the newest model selection, and route both Delete buttons through their visibility managers.
Move Detection to Bboxes onto the shared visibility manager with pre-hiding, load-aware initialization, user-driven updates, and correctly restored workflow state.
Perf
Replace Loader and Smart Detection's fixed initialization/configuration delays with ID-ready fresh-node animation frames and immediate workflow configuration.
Refactor
Resolve the three supported custom LLaVA architectures through lazy named imports instead of a scanner-ambiguous dynamic import while preserving standard Transformers delegation and fallbacks.
2026-09-13
Version: 1.0.18
Fix
Restore older expanded Smart LM Loader workflows that contain both the retired WD14 blacklist and Trust Remote Code values, moving the legacy seed into the current schema order while discarding only obsolete and UI-only entries. Also normalize intermediate saves that contain only the early mode-bar and delete-button placeholders, allowing backend-conditional widgets such as attention mode to resolve correctly.
Version: 1.0.17
Breaking
Remove remote-code options from Smart LM Loader, Registry Manager, model registries, and Docker configuration; models that require repository-supplied Python code are now unsupported.
Fix
Normalize canonical 36-value and frontend-expanded 41-value Smart LM Loader workflows to the 35-value schema before node configuration, discarding only retired or UI-only values while preserving every retained value exactly; also repair interim 40-value saves.
Close and reopen workflows that were already open during the update; refreshing the browser alone may retain the stale in-memory graph.
Explicitly deny remote code at Transformers Auto-class boundaries, keep Florence on its local vendored model and processor implementations, and prevent vLLM plans from emitting the remote-code flag.
Docs
Document the unsupported remote-code boundary and the safe alternatives in the loader, Registry Manager, Docker, and security guides.
2026-09-11
Version: 1.0.16
Fix
Keep Linux Docker auto-start privilege-free by trying only the rootless user service, directing system-service setup to the existing guide instead of invoking or generating privileged commands.
Docs
Remove the recursive Docker/containerd system-data deletion command from the Linux uninstall guide and direct intentional full cleanup to Docker's distribution-specific documentation.
Chore
Remove the built-in software updater and its server endpoints, leaving SmartLLM upgrades to ComfyUI Manager or explicit repository/package-management workflows.
Version: 1.0.15
Fix
Let WD14 honor its established CPU fallback when CUDA is requested but ONNX Runtime does not expose
CUDAExecutionProvider, warning about the effective provider instead of aborting before model load.
Version: 1.0.14
Fix
Remove five Anzhc YOLO artifacts explicitly marked unsafe by Hugging Face from bundled downloads and provenance; upgrade cleanup retires stale remote entries while preserving already-downloaded files as local-only models protected by restricted loading.
Version: 1.0.13
Fix
Keep shipped YOLO registry models selectable before their artifacts exist locally, allowing Smart Detection to invoke the existing pinned, hash-verified automatic download path while continuing to hide missing local-only discoveries.
Show YOLO models in Smart LM Manager with safe add, edit, inspect, download, verify, delete, and registry-removal support, including dedicated checkpoint filename and bbox/segm fields.
Replace vague restricted-loader guidance with the exact supported Ultralytics
>=8.4.67,<8.5and PyTorch>=2.6requirements, detected installed versions, and an interpreter-specific Ultralytics upgrade command.
2026-09-09
Version: 1.0.12
Feat
Turn an empty
user_prompton the duration-neutral MiniMax H3 Scene task into an automatic one-shot story: create an original scene without images, ground it in one image, or connect two ordered endpoints, with synchronized sound, suitable music, and dialogue only when plausible.
Compat
Remove
5sand15sfrom all visible MiniMax H3 task names while preserving retired names as hidden backend and frontend aliases, migrating saved widget values, runtime system prompts, and customized few-shot entries without breaking older workflows or API prompts.Keep explicit timeline generation planned within MiniMax H3's 15-second maximum, and use the same ceiling when automatic story pacing needs a duration bound.
Docs
Document duration-neutral H3 task selection, automatic empty-prompt storytelling, custom-system-prompt behavior, and the distinction between automatic story music and explicit no-music requests.
2026-09-08
Version: 1.0.11
Fix
Apply a shared MiniMax H3 music-intent policy across every backend and task mode so explicit requests for no background music, score, or soundtrack produce
non_diegetic_music: N/Awithout removing scene sound.
Docs
Document the common three-field H3 output contract and the distinction between non-diegetic music and dialogue, ambience, or action sounds.
Version: 1.0.10
Fix
Remove standalone reference-picture alignment metadata from every MiniMax H3 output mode, strip legacy/custom preambles before validation, and keep recovery plus fallback output limited to the three generator-ready H3 fields.
Update bundled H3 prompts and few-shot examples to stop requesting alignment lines, while migrating untouched runtime defaults without replacing user customizations.
Docs
Clarify that MiniMax H3 tasks return only the three required prompt fields.
Version: 1.0.9
Feat
Show the running SmartLLM version in its ComfyUI settings and provide a loopback-only, explicitly confirmed update action that replaces tracked files with official
main, preserves untracked user data, installs requirements, and reports the required restart.
Fix
Allow SmartLLM to start when overlay updates leave inert extracted SmartLLM source files in Eclipse, while preserving Eclipse user-data migration and continuing to block active legacy packs or already-registered
/smartlml/...routes.
Docs
Clarify that startup conflict detection uses active providers and registered routes rather than stale Eclipse filenames.
2026-09-07
Version: 1.0.7
Feat
Replace the generic 15-second MiniMax H3 timeline with explicit T2VA, I2VA, FL2VA, and L2VA tasks, centralized mode metadata, exact reference-count validation, ordered endpoint grouping, natural-language shot counts, requested camera views, and named Tracking Shots.
Expand neutral H3 few-shot coverage in both training profiles for every input mode, including ordered three-view stories and endpoint-convergent motion paths.
Fix
Make explicit FL2VA accept either one first-frame prompt-writing image or an ordered first/last pair, use two-image pose/framing differences to plan subject and camera transitions, keep one-image output grounded in its source, and still land on the downstream encoder-supplied Picture 2.
Require every H3 response to include its complete shot timeline, soundscape, and music fields so source-only or sound-only completions are explicitly invalid.
Upgrade untouched bundled prompt entries individually during default migration while preserving customized values and any inert retired Timeline entry.
Rebuild H3 instructions and both few-shot profiles around Wan-style input/output pairs so models use reference images as visual constraints, focus output on the requested action instead of re-captioning appearance, avoid invented names/entities/story events, and return only plain-text H3 fields without follow-up questions or Markdown sections.
Generalize Wan, H3, and LTX role prompts to the neutral cinematic motion prompt writer occupation, remove model/vendor names from both few-shot profiles' instructional messages.
Docs
Document H3 mode selection, image ordering, shot and camera controls, Tracking Shots, and the separation between SmartLLM prompt-writing references and downstream H3 encoder keyframes.
Breaking
Remove MiniMax H3 Timeline 15s from task registration and bundled defaults without a compatibility alias; select one of the four explicit timeline modes instead.
2026-09-06
Version: 1.0.6
Feat
Add optional-image MiniMax H3 Scene 5s and MiniMax H3 Timeline 15s tasks that turn short stories or numbered shot lists into H3's audio-video prompt structure, including conditional first-frame references, concise image anchors, dialogue, soundscape, music, and duration-safe shot timing.
Add neutral few-shot examples for both MiniMax H3 tasks to the standard and NSFW training profiles so strict field and timeline formatting remains available in either profile.
Docs
Document MiniMax H3 text-to-video, image-to-video, scene, timeline, and Training-chip usage in the Smart LM Loader guide.
2026-08-29
Version: 1.0.4
Feat
Add a qualified Ollama runtime-version selector to Smart LM Manager's Docker Images tab. Installing a selected 0.33.1 or legacy-compatible 0.20.2 image persists its immutable vendor-specific pin and lets the existing container-spec check recreate Ollama on its next start without changing registry entries or deleting model data.
Add per-backend Stop controls for SmartLLM-managed Docker containers. Container shutdown reuses the server-side model-maintenance gate and returns a busy response instead of interrupting an active SmartLLM execution.
Fix
Keep a Transformers VLM on its effective device for every mapped/list item and defer Keep Loaded off cleanup until the complete node execution finishes, preventing later prompts from sending CUDA token IDs into a prematurely CPU-offloaded embedding layer.
Preserve Ollama tensor-shape model-load failures as actionable installed-artifact incompatibility errors, including the rejected tensor and expected/actual dimensions, instead of replacing them with a misleading model-file-not-found diagnosis.
2026-08-28
Version: 1.0.3
Feat
Add a Docker Images view to the Smart LM Manager (Registry Manager) with Docker Engine, daemon, group, and GPU readiness diagnostics; terminal-only Linux installer guidance; vendor-specific managed-image status; and guarded install, update, and removal controls that refuse images still used by containers.
Report Docker image installation and removal milestones in the ComfyUI console, with filtered pull progress and command details available at SmartLLM's debug log level.
Fix
Update the qualified Ollama NVIDIA/CPU and ROCm fallback images from 0.20.2 to 0.33.1 so Qwen 3.8 models use a compatible backend.
Make the standalone Docker image manager pull Ollama's current release channel, detect the installed version and repository digest, and atomically record that immutable pin in SmartLLM's active Docker configuration.
2026-08-22
Version: 1.0.2
Feat
Configurable chip accent: Add a SmartLLM-owned color picker that persists a validated hexadecimal accent in private
config.jsonand applies derived hover, border, trigger, and contrast colors to chip bars and selected chips immediately.Aligned chip popovers: Match Smart LM Loader and Smart Detection popup widths to their rendered chip bars without stretching or shrinking individual chips.
Fix
Set the Smart LM Registry Editor and its model-download surface background to
#3a3a3a.Restore the configured chip-bar surface and interactive popup on Smart LM Loader and Smart Detection by keeping each widget's CSS prefix synchronized with its injected stylesheet in both Nodes 2.0 and classic renderers; also release deferred outside-click listeners whenever the popup closes.
2026-08-19
Version: 1.0.1
Refactor
Move Smart LM Loader and Smart Detection from the Eclipse node menu into the pack-owned
Smart LM Loader → Loadermenu without changing their serialized IDs.Move Detection to Bboxes and its conditional-widget frontend from Eclipse into the pack-owned
Smart LM Loader → Conversionmenu, preserving its data, mask, bbox, list, and workflow contracts.
Fix
Extend the compatibility guard to reject an active Eclipse release that still registers Detection to Bboxes, preventing duplicate node and frontend ownership during partial upgrades.
Docs
Replace the compact landing page with a Nodes 2.0 visual walkthrough covering Smart LM modes, multi-task chaining, WD14 tagging, Smart Detection, detection conversion, and Registry Manager model acquisition.
Clarify task-owned system prompts, optional
user_promptcontext, complete connected-system-prompt overrides, focused-part detection, Wan/LTX image-to-video prompting, and paste-ready song lyric generation.Add the backend-specific copy-paste model registry reference transferred from Eclipse, aligned with the current YOLO
repo_idschema and standard model directories.
2026-08-17
Version: 1.0.0
Feat (New)
Standalone Smart LM Loader and Smart Detection provider preserving the two
[Eclipse]workflow node IDs, schemas, seed behavior, list semantics, and serialized widgets.SmartLLM-owned Registry Manager, settings, frontend helpers, and
smartllm:registry-changedrefresh event.Native, Transformers, GGUF, WD14, YOLO, Florence-2, vLLM, SGLang, Ollama, and llama.cpp infrastructure with verified acquisition and Docker isolation.
Atomic precedence-based migration from Eclipse and legacy SmartLML runtime data without moving model artifacts or transient state.
Fix
Guard every backend-writing setting against ComfyUI's automatic first change callback while hydrating defaults and credential masks only from SmartLLM's redacted endpoint.
Refactor
Present stable
SmartLLM.*controls under the independentSmart LM Loader → Configurationcategory while retainingComfyUI_SmartLLMpackage branding and/smartlmlroutes.Upgrade migration markers atomically with value-free examined-key confirmation while preserving destination config precedence and hiding derived absolute model paths.
Docs
Installation, migration, Registry Manager, Smart LM, Smart Detection, Docker, security, and third-party attribution guides.
Document independent Eclipse, Smart Model Loader, and Smart LM Loader settings and configuration ownership.
ComfyUI SmartLLM
One adaptive interface for vision-language models, text models, WD14 taggers, Florence grounding, and YOLO detection—plus a registry manager that keeps model identity, acquisition, and trust decisions explicit.
Version 1.0.1 provides three Nodes 2.0-ready nodes. ComfyUI Eclipse is optional.

Why use it?
One language-model node adapts to vision, text, and tagger families instead of exposing every backend control at once.
One detection node covers Florence and Qwen grounding tasks alongside YOLO bounding-box and segmentation models.
A matching postprocessor converts structured detections into selectable masks and SAM2-compatible bounding boxes without requiring Eclipse.
Several execution backends share one registry-driven model selector, including Transformers, GGUF, Ollama, vLLM, SGLang, and llama.cpp.
Verified acquisition records model source, immutable revision, integrity, and provenance before committing local files.
Workflow compatibility preserves the historical
[Eclipse]node IDs, so existing workflows load without node replacement.
Visual tour
Start with Smart LM Loader
Search for Smart LM Loader [Eclipse] in Add Node. Choose a registered model, then select the task it should perform. The interface responds to the model family: a vision model exposes image-aware tasks, while text-only and tagger models show their relevant inputs.

The selected task loads its task-specific system prompt automatically, so user_prompt is normally empty for tasks such as Detailed Description. Use it only for additional information, a question, or constraints that are not already expressed by the task. Connecting system_prompt is an explicit override: it switches the node to Direct Chat and bypasses the default task prompt templates. The connected text becomes the complete system instruction, so it must contain every important role, objective, constraint, and output-format requirement. user_prompt then supplies the user message or source material for that custom instruction.
The image input is optional for text-only tasks. Output sockets return generated text plus an image when the selected task supports one.
Some generative tasks intentionally use user_prompt as their source material:
For a Wan or LTX image-to-video task, connect the starting image and describe the intended motion, action, dialogue, style, or camera behavior in
user_prompt. The task prompt tells the model how to format the result; the image establishes visual details such as the person's appearance, while the user text explains what should happen in the video.For Song Lyrics, enter a short story, theme, mood, or song concept in
user_prompt. Genre, language, tempo, or structural preferences can be added when they matter. The result is a structured lyric sheet that can be copied into Suno, Mureka, or another music-generation tool.
These cases do not require a custom system_prompt; selecting the task still loads the appropriate instructions automatically.
Open the mode chip bar
Click the green mode bar to change memory, prompt, advanced-runtime, trust, and model-maintenance behavior. Selected chips are serialized with the workflow; inactive sections stay hidden.

Available Smart LM modes are Cleanup, Keep Loaded, Multi-Task, Training, Advanced, Use Advanced, ⚠ Trust Remote Code, and Delete.
CleanupandKeep Loadedcontrol the model lifecycle between executions.Multi-Taskexposes a sequential task chain;Trainingadds curated task-specific examples to the prompt.Advancedreveals sampling, device, and compile controls.Use Advanceddecides whether the advanced sampling values are applied.⚠ Trust Remote Codepermits pinned repository Python code for the selected model. Enable it only for a source you trust.Deletereveals the separately confirmed local-file deletion action.
Chain tasks in one execution
Enable Multi-Task to expose Task 2 through Task 4. Each active stage receives the previous stage's text and passes its result forward. This is useful for a visual description → prose conversion → prompt refinement sequence without adding several model nodes.

Set unused stages to None. The final active stage becomes the node's text output, while the image output remains available for compatible workflows.
Switch to WD14 tagging
Selecting a WD14 registry entry replaces language-generation widgets with the tagger's general threshold, character threshold, and underscore-formatting controls.

Connect an image and consume the generated tags from the text socket. The image socket passes the source image through for downstream routing.
Run YOLO detection
Search for Smart Detection [Eclipse], select a YOLO registry model, and optionally filter among the classes that specific detector knows. For example, the face-specific face_yolov8m model can use face; other registered models specialize in regions such as eyes, faces, hands, or people. Detector-specific confidence, NMS, filtering, and region-selection controls appear automatically.

The node returns an annotated image, a combined mask, Impact-compatible SEGS, and structured detection data.
Ground phrases and adjust regions
Florence and compatible Qwen models expose language-directed tasks such as Caption to Phrase Grounding. Enter focused parts such as eye;face;mouth; the semicolon-separated phrases are run as individual targets and their regions are merged. Enable Preview Boxes to annotate the result and Adjust to reveal drop-size, crop-factor, and dilation controls.

Smart Detection modes are Cleanup, Keep Loaded, Preview Boxes, Adjust, Advanced, and Delete. The selected model and task determine which controls are meaningful.
Convert detection data to masks and boxes
Connect Smart Detection's image and data outputs to Detection to Bboxes [Eclipse]. The converter accepts regular boxes, OCR quad boxes, and polygons; it can combine them into one mask or return selected regions separately, with optional inversion, grow/shrink, and blur processing. Its bbox output uses the established BBOXES structure for downstream tools such as SAM2 Ultra.
The converter also has an independent image-analysis mode. Enable get_mask_from_image to detect bright or color-channel regions directly with threshold and minimum-area controls instead of consuming JSON data.
Inspect and acquire models
Open SmartLLM → Edit Smart LM Registry (Beta), use the LM Registry left-toolbar launcher, or use the classic-menu button. Registry Manager separates model identity and trust policy from each action you may take.

A registry entry can define its display name, backend, model family, repository or model ID, source, immutable revision, vision capability, local-only policy, remote-code permission, expected SHA-256 digests, and description.
Use Inspect before Save Entry or Download. Verify Local Files does not download anything. Delete Local Files and Remove Registry Entry are separate confirmed operations: removing an entry does not delete its local model files.
Included nodes
NodeInputsOutputsPurposeSmart LM Loader [Eclipse]optional images, optional system prompt, adaptive widgetsimage, textRun registered vision-language, text, or WD14 modelsSmart Detection [Eclipse]image, adaptive widgetsimage, mask, SEGS, dataRun Florence/Qwen grounding or YOLO detectionDetection to Bboxes [Eclipse]image, optional detection data, mask controlsmask, BBOXESConvert Smart Detection data or image regions into masks and boxes
The [Eclipse] suffixes are compatibility identifiers. SmartLLM owns all three implementations and does not require Eclipse at runtime. The two model nodes appear under Smart LM Loader → Loader, while the converter appears under Smart LM Loader → Conversion. Their historical IDs remain unchanged so saved workflows continue to resolve without node replacement.
Supported backends and model families
PathTypical useTransformersLocal Hugging Face vision-language and text models, including Qwen, Florence, Mistral, and LLaVA-family modelsGGUF / llama.cppQuantized local language and vision-language executionOllamaModels served by an Ollama runtimevLLM / native vLLMHigh-throughput model serving, locally or through the managed Docker pathSGLangStructured high-throughput model servingWD14Image tagging with independent general and character thresholdsYOLOBounding-box and segmentation detection
Backend availability depends on the optional packages or services installed in your ComfyUI environment. The base installation does not force every compiled runtime onto the user.

