Commit 4e154b5
fix(ci): unbreak rerankers (torch bump) and vllm-omni on aarch64 (#9688)
Two unrelated CI breakages bundled together since both are one-liners:
- rerankers: bump torch 2.4.1 -> 2.7.1 on cpu/cublas12. The unpinned
transformers resolves to 5.x, whose moe.py registers a custom_op with
string-typed `'torch.Tensor'` annotations that torch 2.4.1's
infer_schema rejects, blocking the gRPC server from starting and
failing all 5 backend tests with "Connection refused" on :50051.
Matches the version used by the transformers backend.
- vllm-omni: strip fa3-fwd from the upstream requirements/cuda.txt
before resolving on aarch64. fa3-fwd 0.0.3 ships only an
x86_64 wheel and has no sdist, making the cuda profile unsatisfiable
on Jetson/SBSA. fa3-fwd is a soft runtime dep — vllm-omni's
attention backends fall back to FA2 then SDPA when it's missing.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>1 parent 969005b commit 4e154b5
3 files changed
Lines changed: 10 additions & 2 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
1 | 1 | | |
2 | 2 | | |
3 | | - | |
| 3 | + | |
4 | 4 | | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
1 | 1 | | |
2 | 2 | | |
3 | | - | |
| 3 | + | |
4 | 4 | | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
79 | 79 | | |
80 | 80 | | |
81 | 81 | | |
| 82 | + | |
| 83 | + | |
| 84 | + | |
| 85 | + | |
| 86 | + | |
| 87 | + | |
| 88 | + | |
| 89 | + | |
82 | 90 | | |
83 | 91 | | |
84 | 92 | | |
| |||
0 commit comments