You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Address second review round: symmetric base_layer lookup, loud miss, value oracle
- _restore_qtensor_wrappers: normalize the .base_layer suffix on both the saved
q_tensor_state keys and the module names, so the lookup works whether the model
was compressed before adapters were attached (quantize.py --compress) or after
(QATTrainer._quantize_model). Warn when saved weights match no module at all,
which is the silent failure this PR set out to fix.
- export.py: document why skipping the restore is safe (QATTrainer snapshots the
state at trainer init from the base state from_pretrained already restored).
- LoRA-QAT example test: compare exported scales against a direct PTQ export
rather than only asserting key presence.
- Unit tests: cover the reverse key direction and the no-match warning.
- CHANGELOG: add the 0.47 Bug Fixes entry.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
Copy file name to clipboardExpand all lines: CHANGELOG.rst
+2Lines changed: 2 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -21,6 +21,8 @@ Changelog
21
21
22
22
**Bug Fixes**
23
23
24
+
- Fix QLoRA export in ``examples/llm_qat/export.py`` failing with ``AssertionError: Model already has modelopt state!`` (NVBug 6542481). The QLoRA training output is an adapter-only checkpoint, so ``from_pretrained`` resolves the quantized base model from ``adapter_config.json`` and ``enable_huggingface_checkpointing`` already restores its ModelOpt state; the export then restored a second time. It now restores only when the loaded model is not already converted. Two further breakages on the same path are also fixed: ``_restore_qtensor_wrappers`` matched no modules because PEFT re-parents the quantized linear as ``<name>.base_layer`` while ``q_tensor_state`` is keyed by the name it was saved with (the packed NVFP4 weight then reached ``F.linear`` and raised a shape error), and ``postprocess_state_dict`` silently dropped every ``base_layer.*`` key missing from a hand-maintained rename map — losing the NVFP4 ``weight_scale_2`` global scale and any linear ``bias`` (Qwen2-style q/k/v biases), and leaving ``base_layer`` in the exported AWQ ``pre_quant_scale`` key. The rename is now a generic ``.base_layer.`` strip.
0 commit comments