#9585·InvokeAI

[bug]: Crashing Krea2 non-standard checkpoints after recent update.

Author: CarstenD74Created Sep 13, 2026Updated Sep 14, 2026
Labelsbug
### Is there an existing issue for this problem? - [x] I have searched the existing issues ### Install method Docker image on unRAID server ### Operating system Linux (unRAID) ### GPU vendor AMD (ROCm) ### GPU model AI PRO R9700 ### GPU VRAM 32 GiB ### Version number v6.14.1-post1 ### Browser Edge Version 150.0.4078.48 (Official build) (64-bit) ### System Information
  |   -- | -- version | "6.14.0" dependencies |   absl-py | "2.4.0" accelerate | "1.14.0" annotated-doc | "0.0.4" annotated-types | "0.7.0" anyio | "4.14.1" argon2-cffi | "25.1.0" argon2-cffi-bindings | "25.1.0" arrow | "1.4.0" asttokens | "3.0.1" async-lru | "2.3.0" attrs | "26.1.0" babel | "2.18.0" bcrypt | "3.2.2" beautifulsoup4 | "4.15.0" bidict | "0.23.1" bitsandbytes | "0.49.2" blake3 | "1.0.9" bleach | "6.4.0" certifi | "2026.6.17" cffi | "2.0.0" charset-normalizer | "3.4.7" click | "8.4.2" coloredlogs | "15.0.1" comm | "0.2.3" compel | "2.4.0" contourpy | "1.3.3" cryptography | "49.0.0" CUDA | "N/A" cycler | "0.12.1" debugpy | "1.8.21" decorator | "5.3.1" defusedxml | "0.7.1" Deprecated | "1.3.1" diffusers | "0.39.0" dnspython | "2.8.0" dynamicprompts | "0.31.0" ecdsa | "0.19.2" einops | "0.8.2" email-validator | "2.3.0" executing | "2.2.1" fastapi | "0.141.1" fastapi-events | "0.12.2" fastjsonschema | "2.21.2" filelock | "3.29.4" flatbuffers | "25.12.19" fonttools | "4.63.0" fqdn | "1.5.1" fsspec | "2026.6.0" gguf | "0.19.0" h11 | "0.16.0" hf-xet | "1.5.1" httpcore | "1.0.9" httptools | "0.8.0" httpx | "0.28.1" huggingface_hub | "1.21.0" humanfriendly | "10.0" idna | "3.18" ImageIO | "2.37.4" imageio-ffmpeg | "0.6.0" importlib_metadata | "9.0.0" InvokeAI | "6.14.0" ipykernel | "7.3.0" ipython | "9.15.0" ipython_pygments_lexers | "1.1.1" isoduration | "20.11.0" jax | "0.7.1" jaxlib | "0.7.1" jedi | "0.20.0" Jinja2 | "3.1.6" json5 | "0.15.0" jsonpointer | "3.1.1" jsonschema | "4.26.0" jsonschema-specifications | "2025.9.1" jupyter-events | "0.12.1" jupyter-lsp | "2.3.1" jupyter_builder | "1.0.2" jupyter_client | "8.9.1" jupyter_core | "5.9.1" jupyter_server | "2.20.0" jupyter_server_terminals | "0.5.4" jupyterlab | "4.6.0" jupyterlab_pygments | "0.3.0" jupyterlab_server | "2.28.0" kiwisolver | "1.5.0" lark | "1.3.1" markdown-it-py | "4.2.0" MarkupSafe | "3.0.3" matplotlib | "3.11.0" matplotlib-inline | "0.2.2" mdurl | "0.1.2" mediapipe | "0.10.14" mistral_common | "1.11.6" mistune | "3.3.2" ml_dtypes | "0.5.4" mpmath | "1.3.0" nbclient | "0.11.0" nbconvert | "7.17.1" nbformat | "5.10.4" nest-asyncio2 | "1.7.2" networkx | "3.6.1" notebook | "7.6.0" notebook_shim | "0.2.4" numpy | "1.26.4" onnx | "1.16.1" onnxruntime | "1.19.2" opencv-contrib-python | "4.11.0.86" opt_einsum | "3.4.0" packaging | "26.2" pandocfilters | "1.5.1" parso | "0.8.7" passlib | "1.7.4" pexpect | "4.9.0" picklescan | "1.0.4" pillow | "12.2.0" platformdirs | "4.10.0" prometheus_client | "0.25.0" prompt_toolkit | "3.0.52" protobuf | "4.25.9" psutil | "7.2.2" ptyprocess | "0.7.0" pure_eval | "0.2.3" pyasn1 | "0.6.3" pycountry | "26.2.16" pycparser | "3.0" pydantic | "2.13.4" pydantic-extra-types | "2.11.1" pydantic-settings | "2.14.2" pydantic_core | "2.46.4" Pygments | "2.20.0" pyparsing | "3.3.2" PyPatchMatch | "1.0.2" python-dateutil | "2.9.0.post0" python-dotenv | "1.2.2" python-engineio | "4.13.3" python-jose | "3.5.0" python-json-logger | "4.1.0" python-multipart | "0.0.32" python-socketio | "5.16.3" PyWavelets | "1.9.0" PyYAML | "6.0.3" pyzmq | "27.1.0" referencing | "0.37.0" regex | "2026.5.9" requests | "2.34.2" rfc3339-validator | "0.1.4" rfc3986-validator | "0.1.1" rfc3987-syntax | "1.1.0" rich | "15.0.0" rpds-py | "2026.5.1" rsa | "4.9.1" safetensors | "0.8.0" scipy | "1.17.1" semver | "3.0.4" Send2Trash | "2.1.0" sentencepiece | "0.2.0" setuptools | "82.0.1" shellingham | "1.5.4" simple-websocket | "1.1.0" six | "1.17.0" sounddevice | "0.5.5" soupsieve | "2.8.4" spandrel | "0.4.2" stack-data | "0.6.3" starlette | "0.48.0" sympy | "1.14.0" terminado | "0.18.1" tiktoken | "0.13.0" tinycss2 | "1.5.1" tokenizers | "0.22.2" torch | "2.10.0+rocm7.1" torchsde | "0.2.6" torchvision | "0.25.0+rocm7.1" tornado | "6.5.7" tqdm | "4.68.3" traitlets | "5.15.1" trampoline | "0.1.2" transformers | "5.5.4" triton-rocm | "3.6.0" typer | "0.25.1" typing-inspection | "0.4.2" typing_extensions | "4.15.0" tzdata | "2026.2" uri-template | "1.3.0" urllib3 | "2.7.0" uvicorn | "0.49.0" uvloop | "0.22.1" watchfiles | "1.2.0" wcwidth | "0.8.1" webcolors | "25.10.0" webencodings | "0.5.1" websocket-client | "1.9.0" websockets | "16.0" wrapt | "2.2.2" wsproto | "1.3.2" zipp | "4.1.0" config |   schema_version | "4.0.3" legacy_models_yaml_path | null host | "0.0.0.0" port | 9090 allow_origins | [] allow_credentials | true allow_methods |   0 | "*" allow_headers |   0 | "*" ssl_certfile | null ssl_keyfile | null base_url | null forwarded_allow_ips | "127.0.0.1" http_compression_level | 9 log_tokenization | false patchmatch | true models_dir | "models" convert_cache_dir | "models/.convert_cache" download_cache_dir | "models/.download_cache" legacy_conf_dir | "configs" db_dir | "databases" outputs_dir | "outputs" image_subfolder_strategy | "flat" custom_nodes_dir | "nodes" style_presets_dir | "style_presets" workflow_thumbnails_dir | "workflow_thumbnails" log_handlers |   0 | "console" log_format | "color" log_level | "info" log_sql | false log_level_network | "warning" use_memory_db | false dev_reload | false profile_graphs | false profile_prefix | null profiles_dir | "profiles" max_cache_ram_gb | 22 max_cache_vram_gb | null log_memory_usage | false model_cache_keep_alive_min | 0 device_working_mem_gb | 3 enable_partial_loading | false keep_ram_copy_of_weights | false ram | null vram | null lazy_offload | true pytorch_cuda_alloc_conf | null device | "auto" generation_devices | "auto" offload_text_encoders_to_idle_gpus | true precision | "bfloat16" sequential_guidance | false wan_memory_optimization | false pid_memory_optimization | false attention_type | "torch-sdp" attention_slice_size | "auto" force_tiled_decode | true pil_compress_level | 1 max_queue_size | 10000 session_queue_mode | "round_robin" clear_queue_on_startup | true max_queue_history | null allow_nodes | null deny_nodes | null node_cache_size | 512 hashing_algorithm | "blake3_single" remote_api_tokens | null scan_models_on_startup | false allow_private_download_urls | false download_proxy | null unsafe_disable_picklescan | false allow_unknown_models | true multiuser | false strict_password_checking | false external_alibabacloud_api_key | null external_alibabacloud_base_url | null external_gemini_api_key | null external_openai_api_key | null external_gemini_base_url | null external_openai_base_url | null external_seedream_api_key | null external_seedream_base_url | null set_config_fields |   0 | "max_cache_ram_gb" 1 | "precision" 2 | "attention_type" 3 | "host" 4 | "force_tiled_decode" 5 | "legacy_models_yaml_path" 6 | "port" 7 | "keep_ram_copy_of_weights" 8 | "enable_partial_loading" 9 | "clear_queue_on_startup"
### What happened

[bug]: FP8 and other quantized checkpoints are expanded to full precision on load — resident size is identical regardless of quantization level (ROCm)

Is there an existing issue for this problem?

  • [x] I have searched the existing issues

Operating system

Linux (unRAID 7.x, Docker)

GPU vendor

AMD (ROCm)

GPU model

<!-- FILL IN: exact card, gfx1201 / RDNA 4 -->

GPU VRAM

32 GB

Version number

6.14.0 and 6.14.1-post1 — identical behaviour on both.

  • ghcr.io/invoke-ai/invokeai:6.14.0-rocmsha256:032161d7873e3d3f13f699fa9a8a10ec3f5a3e86b662f07843f263fd1a245580
  • ghcr.io/invoke-ai/invokeai:main-rocm (6.14.1-post1) — sha256:2bb1a68525a8c6e567ee499ac391a8289c469414fc9b9bea74eeb37caa7d70de

Browser

<!-- FILL IN -->

What happened

On ROCm, quantized checkpoints appear to be dequantized to full precision during loading. The memory saving that quantization is chosen for is lost entirely, and the resident footprint is the same no matter which quantization is used.

GGUF checkpoints are unaffected — they load at approximately their on-disk size.

Evidence 1: three different Z Image quantizations, one resident size

Three distinct Z Image checkpoints, three different file sizes, all reporting the same loaded size:

Model ID | Size on disk | Total model size as loaded -- | -- | -- d9006a28-cc54-49d3-93d9-3e1652d2cdf5 | 6.16 GiB | 11,739.56 MB 00ce5266-f09d-4287-a4ea-d8ae89f56263 | 6.54 GiB | 11,739.56 MB f50e705c-df8a-496f-9acf-491fba3d885f | 12.57 GiB | 11,739.56 MB

Two independent FP8 Krea-2 checkpoints both land at exactly 24,452.35 MB. The GGUF build of the same model family stays at its on-disk size.

precision: bfloat16 is set in invokeai.yaml, and ~2x is what FP8 → bf16 expansion would produce.

Consequence

The model no longer fits in the cache, so it is evicted and re-staged on every generation:

  • GGUF Krea-2 (12.69 GiB resident, ~21 GiB working set with the encoder): stays cached, repeated in 0.00s cache hits, 22 s per 960x1344 image.
  • FP8 Krea-2 (23.88 GiB resident, ~32 GiB working set): never a single cache hit, full re-stage every generation, 82 s per image, and OOM-killed outright at container limits below 26 GiB.

The staging phase is host-RAM bound with the GPU idle

Sampling the container cgroup's memory.current and amdgpu's mem_info_vram_used every 2 seconds during an FP8 load:

13:58:57  sys=  8.2 GiB  vram=  0.6 GiB    <- load begins
13:59:07  sys= 20.3 GiB  vram=  0.6 GiB
13:59:19  sys= 26.0 GiB  vram=  0.6 GiB    <- pinned at container limit
   ... 45 seconds pinned, GPU idle at 0.6 GiB ...
14:00:06  sys= 25.6 GiB  vram=  1.2 GiB    <- transfer to VRAM begins
14:00:12  sys=  6.2 GiB  vram= 20.6 GiB
14:00:16  sys=  2.5 GiB  vram= 26.0 GiB    <- steady state

The weights are fully materialised in system RAM, then copied to VRAM in ~6 seconds. Steady-state host RAM afterwards is 2.5 GiB. So host RAM must be at least model-sized during loading even when VRAM is abundant, and the expansion above doubles what that costs.

What you expected to happen

That an FP8 checkpoint would occupy roughly its on-disk size in memory, and that selecting a smaller quantization would reduce the memory footprint.

How to reproduce the problem

  1. ROCm container image on a system with 32 GB VRAM.
  2. Install two quantizations of the same model — e.g. a 6 GiB and a 12 GiB Z Image checkpoint, or an FP8 and a Q8_0 GGUF Krea-2.
  3. Load each and compare the Total model size value in the [MODEL CACHE] Loaded model log line against the file size on disk.
  4. The quantized non-GGUF checkpoints report a resident size unrelated to their file size.

Possible cause

The ROCm image does not contain rocminfo, so bitsandbytes cannot detect the GPU architecture and falls back. Logged at every startup:

Could not detect ROCm GPU architecture: [Errno 2] No such file or directory: 'rocminfo'
ROCm GPU architecture detection failed despite ROCm being available.
Could not detect ROCm warp size: [Errno 2] No such file or directory: 'rocminfo'.
Defaulting to 64. (some 4-bit functions may not work!)

If quantized kernels are unavailable because the architecture could not be identified, dequantizing the weights at load would be a plausible fallback — and would explain why GGUF (which has its own dequantization path) behaves differently.

This is speculation on my part; I have not instrumented it. But rocminfo being absent from the official ROCm image looks like a packaging defect regardless.


Environment

MALLOC_MMAP_THRESHOLD_=1048576
DISABLE_PINNED_MEMORY=1
PYTORCH_HIP_ALLOC_CONF=garbage_collection_threshold:0.8,max_split_size_mb:128
HSA_OVERRIDE_GFX_VERSION=12.0.1
HSA_ENABLE_SDMA=0
PYTORCH_TUNABLEOP_ENABLED=0
TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=0
MIOPEN_FIND_MODE=FAST
GPU_DRIVER=rocm
yaml
schema_version: 4.0.3
max_cache_ram_gb: 14.0
keep_ram_copy_of_weights: false
device_working_mem_gb: 3.0
model_cache_keep_alive_min: 0
enable_partial_loading: false
force_tiled_decode: true
device: auto
precision: bfloat16
attention_type: torch-sdp

Host: 32 GB system RAM (31 GiB usable), AMD Ryzen 7 3800X (8C/16T), container limited to 26 GiB and 6 CPUs. Nothing else memory-heavy running. Components in the Krea-2 workflow: Qwen3-VL 4B encoder (8,464.46 MB), Qwen Image VAE (242.03 MB).

What this is not

Recorded so others don't repeat the dead ends I worked through:

  • Not a leak. Host RAM returns to 2.5–7.8 GiB after every load; nothing accumulates across generations.
  • Not a regression. 6.14.0 and 6.14.1-post1 are identical. Earlier versions could not be tested — Krea-2 support does not exist in 6.13.x.
  • Not fixable via max_cache_ram_gb. Tested at 8.0, 12.0 and 14.0.
  • Not glibc fragmentation. The FAQ's MALLOC_MMAP_THRESHOLD_=1048576 workaround was in place throughout, with DISABLE_PINNED_MEMORY=1.
  • Not a GPU fault. No amdgpu ring timeouts, resets or segfaults in dmesg. All OOM kills are CONSTRAINT_MEMCG, contained to the container cgroup, with no Python traceback.

Three smaller issues noticed along the way

Happy to split these out.

1. The RAM cache statistics block never updates. Byte-identical across consecutive generations (hits: 2 / misses: 4 / cached: 1 / cleared: 1) despite each run loading several models.

2. Model load time is attributed to the denoise node. krea2_denoise was billed at 65.2 s while the 9 sampling steps took ~20 s (2.26 s/it) and the Loaded model line reported 3.16 s. The ~45 s of staging is folded into denoise and invisible in the graph stats — the timing table points at sampling when the cost is in loading. This made the problem substantially harder to locate.

3. force_tiled_decode: true appeared to have no effect. VRAM still spiked to 31.6 GiB at VAE decode with it enabled. Possibly not wired to the Qwen Image VAE path.

### What you expected to happen I expected to be able to run the models. ### How to reproduce the problem Start the app load any model which is not quantized. ### Additional context _No response_ ### Discord username C_Dollerup