Pinned one-DGX-Spark Docker recipe for DeepSeek V4 Flash with EXL3, SparkInfer, and 262K NVFP4 MLA K
reasoning_effort=max consumes entire output budget at ~289k context, 0 visible tokens (finish_reason=length)
No comments yet. Be the first to share.