Skip to content

Qwen3.6-27B recipe and README optimized for agentic coding #recipebot#267

Draft
jawadaminGOOG wants to merge 1 commit into
AI-Hypercomputer:mainfrom
jawadaminGOOG:qwen-recipe
Draft

Qwen3.6-27B recipe and README optimized for agentic coding #recipebot#267
jawadaminGOOG wants to merge 1 commit into
AI-Hypercomputer:mainfrom
jawadaminGOOG:qwen-recipe

Conversation

@jawadaminGOOG

Copy link
Copy Markdown

This PR contains modular recipes to deploy, optimize, and integrate the Qwen/Qwen3.6-27B-FP8 model on Google Kubernetes Engine (GKE) targeting Cloud TPU v6e-4 (4 chips, 2x2 topology) with vLLM.

UBENCH Results (Cloud TPU v6e-4):

  • uBench Run ID: vllm_inference-qwen3_32b-fp8-2026-07-13_194548-04391392-a560-424f-a308-e71e591ac9c9
  • Jobset Name: jawadamin-ubench-dw8wrtyo
  • Model: Qwen/Qwen3.6-27B-FP8
  • Request Throughput: 28.07 req/s
  • TTFT P50: 4529.76 ms
  • Output Token Throughput: 3233.91 tokens/s (404.24 tokens/s/chip)

- Keep final deployment manifests in gke/ and standard 5-step contribution-ready README.
- Include official uBench benchmark results and standard Serving Benchmark Result table.
- Remove experimental subfolders (moved to standalone repository).

UBENCH Results (Cloud TPU v6e-4):
- uBench Run ID: vllm_inference-qwen3_32b-fp8-2026-07-13_194548-04391392-a560-424f-a308-e71e591ac9c9
- Jobset Name: jawadamin-ubench-dw8wrtyo
- Model: Qwen/Qwen3.6-27B-FP8
- Request Throughput: 28.07 req/s
- TTFT P50: 4529.76 ms
- Output Token Throughput: 3233.91 tokens/s (404.24 tokens/s/chip)

TAG=agy
CONV=a3f7389e-c990-4e97-b884-73a3e79ea117
@jawadaminGOOG jawadaminGOOG changed the title Qwen3.6-27B recipe and README optimized for agentic coding Qwen3.6-27B recipe and README optimized for agentic coding #recipebot Jul 13, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant