Skip to content

Commit 6276706

Browse files
authored
kleidiai : fix MUL_MAT support for batched (3D) inputs (#20620)
* kleidiai : fix MUL_MAT support for batched (3D) inputs The supports_op() check incorrectly rejected MUL_MAT operations with 3D inputs (ne[2] > 1), but the actual compute_forward_qx() implementation handles batched inputs correctly via a loop over ne12. This caused models with Q4_0/Q8_0 weights to crash during graph scheduling when n_seq_max > 1, because weights were placed in KLEIDIAI buffers during loading (tested with 2D inputs) but the runtime used 3D inputs. Also relax the buffer check to allow supports_op() to be called during weight loading when src[0]->buffer is NULL. Fixes #20608 * Kleidiai support_ops should only return true for 3D inputs, not also 4D
1 parent 740a447 commit 6276706

1 file changed

Lines changed: 1 addition & 1 deletion

File tree

ggml/src/ggml-cpu/kleidiai/kleidiai.cpp

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1461,7 +1461,7 @@ class extra_buffer_type : ggml::cpu::extra_buffer_type {
14611461
return false;
14621462
}
14631463
if ((op->src[1]->type == GGML_TYPE_F32 || op->src[1]->type == GGML_TYPE_I32) &&
1464-
ggml_ne(op->src[1], 2) == 1 && ggml_ne(op->src[1], 3) == 1) {
1464+
ggml_ne(op->src[1], 3) == 1) {
14651465
return true;
14661466
}
14671467
}

0 commit comments

Comments
 (0)