Add global scale support to quantized layers by aleroot · Pull Request #426 · ml-explore/mlx-swift

aleroot · 2026-06-19T15:58:30Z

Store optional globalScale on QuantizedLinear and QuantizedEmbedding so nvfp4 weights can preserve the scale needed by lower-level MLX quantize/dequantize operations.

Forward the scale when creating or dequantizing weights, add a direct pre-quantized QuantizedEmbedding initializer to match QuantizedLinear, and guard global-scale execution paths on Metal because MLX does not support globalScale dequantization there.

Add focused tests for the new layer state and parameter exposure without broadening the generic quantization API surface.

Proposed changes

Please include a description of the problem or feature this PR is addressing. If there is a corresponding issue, include the issue #.

Checklist

Put an x in the boxes that apply.

I have read the CONTRIBUTING document
I have run pre-commit run --all-files to format my code / installed pre-commit prior to committing changes
I have added tests that prove my fix is effective or that my feature works
I have updated the necessary documentation (if needed)

davidkoski · 2026-06-30T17:41:21Z

    public let mode: QuantizationMode
    public let scales: MLXArray
    public let biases: MLXArray?
+    public let globalScale: MLXArray?


One thing we need to be careful of here is that there is currently no python support for this. That means:

maybe we want to set the key as global_scale or global_scales to match the likely python naming

if this is non-nil after quantizing the model, it will be required when loading weights

will this cause a problem or is nvfp4 just not used?

Good point. I changed the serialised parameter key to global_scale

davidkoski

I like it, but see my questions -- I think these need some thought before merging.

Store optional globalScale on QuantizedLinear and QuantizedEmbedding so nvfp4 weights can preserve the scale needed by lower-level MLX quantize/dequantize operations. Forward the scale when creating or dequantizing weights, add a direct pre-quantized QuantizedEmbedding initializer to match QuantizedLinear, and guard global-scale execution paths on Metal because MLX does not support globalScale dequantization there. Add focused tests for the new layer state and parameter exposure without broadening the generic quantization API surface.

davidkoski reviewed Jun 30, 2026

View reviewed changes

Comment thread Source/MLXNN/Quantized.swift Outdated

davidkoski reviewed Jun 30, 2026

View reviewed changes

davidkoski requested changes Jun 30, 2026

View reviewed changes

aleroot force-pushed the quant_scale branch from 13a6ad5 to 7db225c Compare June 30, 2026 17:56

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Add global scale support to quantized layers#426

Add global scale support to quantized layers#426
aleroot wants to merge 1 commit into
ml-explore:mainfrom
aleroot:quant_scale

aleroot commented Jun 19, 2026

Uh oh!

Uh oh!

davidkoski Jun 30, 2026

Uh oh!

aleroot Jun 30, 2026

Uh oh!

davidkoski left a comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Uh oh!

Conversation

aleroot commented Jun 19, 2026

Proposed changes

Checklist

Uh oh!

Uh oh!

davidkoski Jun 30, 2026

Choose a reason for hiding this comment

Uh oh!

aleroot Jun 30, 2026

Choose a reason for hiding this comment

Uh oh!

davidkoski left a comment

Choose a reason for hiding this comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants