Gittensor Model Hub
We build fast, accurate, and useful AI models optimized for
RTX 5090 and RTX PRO 6000 — making AI more accessible
while remaining powerful and efficient.
Our models are paired with
SparkInfer,
a Blackwell-native, zero-dependency custom inference engine. We intentionally maintain a small
selection of high-quality SOTA models — already achieving
2–3× faster inference than llama.cpp in many workloads.
SparkInfer
Qwen3.8-27B-NVFP4-RTX5090
Spark-Hermes-3.8-27B
Featured Models
An NVIDIA ModelOpt NVFP4-optimized version of Qwen3.8-27B, tuned specifically
for GeForce RTX 5090.
- Native 256K context on 32 GB VRAM
- Faster than Unsloth NVFP4
- Accuracy largely preserved
- Optimized for high-throughput local inference
A co-trained agent model combining Hermes and Qwen3.8-27B,
designed for tool use, structured reasoning, and agentic workflows.
- Model card available first
- Weights coming soon
- Built for practical agent deployments
SparkInfer
Blackwell-native. Zero dependencies. Built for speed.
SparkInfer
is our custom inference engine designed to maximize performance on modern NVIDIA GPUs,
especially RTX 5090 and RTX PRO 6000 systems.
- Custom Blackwell-optimized runtime
- Minimal setup friction
- Focused on real-world inference speed
- Designed to run selected SOTA models faster than common open-source inference stacks
Why Gittensor?
- Fewer models, higher quality
- Optimized for consumer and professional RTX GPUs
- Fast inference without unnecessary dependencies
- Practical models for agents, coding, and general use
- Performance-first engineering
Short version
Gittensor Model Hub builds fast, accurate AI models optimized for RTX 5090 and
RTX PRO 6000. Paired with
SparkInfer,
our Blackwell-native inference engine, our curated SOTA models run
2–3× faster than llama.cpp while remaining practical and efficient.