Popular repositories Loading
-
-
-
-
ninfer-3090
ninfer-3090 PublicForked from Don-Chad/ninfer-3090
Fast Qwen3.8-27B inference on one RTX 3090: ReplaySSM, MTP3, reasoning effort, C1-C8 batching, and native Windows and Linux builds.
C++
-
qwen38-27b-rtx3090
qwen38-27b-rtx3090 PublicForked from syv-ai/qwen38-27b-rtx3090
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-ou…
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

