Problem Description
consolidate register_alg and register_pipeline
gradient accumulate step warning in cli
pass input_id to quantize
fix quantize_layer attention loss
verify algs name
/models/Qwen3-8B --tasks lambada_openai --algs haha --seqlen 512 --iters 1 --gradient_accumulate_steps 2
issue repeated log
/home/wenhuach/miniforge3/envs/autoround/bin/python /home/wenhuach/auto-round/auto_round/main.py /models/Qwen3-0.6B --tasks lambada_openai --algs awq,auto-round --seqlen 512
2026-07-10 17:01:26 INFO main.py L292: start to quantize /models/Qwen3-0.6B
2026-07-10 17:01:26 INFO config.py L53: enable_opt_rtn is turned on, set --disable_opt_rtn for higher speed at the cost of accuracy.
2026-07-10 17:01:26 INFO config.py L53: enable_opt_rtn is turned on, set --disable_opt_rtn for higher speed at the cost of accuracy.
issue no loss print
2026-07-10 17:02:13 INFO mappings.py L560: AWQ resolved 84 smooth-balance mappings (3 per block × 28 blocks).
2026-07-10 17:02:13 INFO mappings.py L423: Using registered AWQ mappings for Qwen3ForCausalLM.
2026-07-10 17:02:13 INFO mappings.py L560: AWQ resolved 84 smooth-balance mappings (3 per block × 28 blocks).
2026-07-10 17:02:13 INFO base.py L201: AWQ: resolved 84 mappings across 28 blocks.
Quantizing model.layers.0: 0%| | 0/28 [00:00<?, ?it/s]2026-07-10 17:02:52 INFO device.py L1450: 'peak_ram': 2.54GB, 'peak_vram': 1.25GB
Quantizing model.layers.1: 4%|▌ | 1/28 [00:39<17:35, 39.10s/it]2026-07-10 17:03:24 INFO device.py L1450: 'peak_ram': 2.57GB, 'peak_vram': 1.37GB
Quantizing model.layers.2: 7%|█▏ | 2/28 [01:10<15:04, 34.80s/it]
it may be better just use inherit instead of mixin for mllm and dissfusion
support mxfp4/nvfp4 opt-rtn
unrecognized keys ['gradient_accumulate_steps'] were passed. Please check them. If you use old api, just ignore this warning
seqlen could not be 1
Reproduction Steps
~
Environment Information
No response
Error Logs
Additional Context
No response
Problem Description
consolidate register_alg and register_pipeline
gradient accumulate step warning in cli
pass input_id to quantize
fix quantize_layer attention loss
verify algs name
/models/Qwen3-8B --tasks lambada_openai --algs haha --seqlen 512 --iters 1 --gradient_accumulate_steps 2
issue repeated log
/home/wenhuach/miniforge3/envs/autoround/bin/python /home/wenhuach/auto-round/auto_round/main.py /models/Qwen3-0.6B --tasks lambada_openai --algs awq,auto-round --seqlen 512
2026-07-10 17:01:26 INFO main.py L292: start to quantize /models/Qwen3-0.6B
2026-07-10 17:01:26 INFO config.py L53: enable_opt_rtn is turned on, set --disable_opt_rtn for higher speed at the cost of accuracy.
2026-07-10 17:01:26 INFO config.py L53: enable_opt_rtn is turned on, set --disable_opt_rtn for higher speed at the cost of accuracy.
issue no loss print
2026-07-10 17:02:13 INFO mappings.py L560: AWQ resolved 84 smooth-balance mappings (3 per block × 28 blocks).
2026-07-10 17:02:13 INFO mappings.py L423: Using registered AWQ mappings for Qwen3ForCausalLM.
2026-07-10 17:02:13 INFO mappings.py L560: AWQ resolved 84 smooth-balance mappings (3 per block × 28 blocks).
2026-07-10 17:02:13 INFO base.py L201: AWQ: resolved 84 mappings across 28 blocks.
Quantizing model.layers.0: 0%| | 0/28 [00:00<?, ?it/s]2026-07-10 17:02:52 INFO device.py L1450: 'peak_ram': 2.54GB, 'peak_vram': 1.25GB
Quantizing model.layers.1: 4%|▌ | 1/28 [00:39<17:35, 39.10s/it]2026-07-10 17:03:24 INFO device.py L1450: 'peak_ram': 2.57GB, 'peak_vram': 1.37GB
Quantizing model.layers.2: 7%|█▏ | 2/28 [01:10<15:04, 34.80s/it]
it may be better just use inherit instead of mixin for mllm and dissfusion
support mxfp4/nvfp4 opt-rtn
unrecognized keys ['gradient_accumulate_steps'] were passed. Please check them. If you use old api, just ignore this warning
seqlen could not be 1
Reproduction Steps
~
Environment Information
No response
Error Logs
Additional Context
No response