GENESIS ARCHIVE + COMMUNITY · SEPTEMBER 2026

What people do
for real with local.

A living catalog of local LLM setups, connected to their hardware, stack and results. Benchmarks tell you what fits. This tells you what works.

Born from a r/LocalLLM thread · September 2026

155 Genesis records1 community snapshotsOur method

This wiki is only as useful as the details people share. Take your time — every careful contribution makes the map clearer for everyone.

Sort by
Compare
Capacity range
Exact reported VRAM
Practice
Model
Quantization
Runtime
Tools / workflow (harness)
Record origin
EXPLORE BY CAPACITY

Start with your machine.

Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.

156 resultsEvery record keeps its origin visible.
Hybrid orchestrationCommunity snapshot2026-09
27B reasoning model on 2x RTX 2080 Ti (NVFP4, vLLM) for agentic orchestration

Qwen3.8-27B (NVFP4 W4A16 + FP8 KV) served on 2x RTX 2080 Ti 22GB over NVLink via a vLLM TP2 fork. 196K context cap, MTP K=3, prefix caching on. Feeds a 1-orchestrator + 7-worker harness: ~1227 tok/s prefill and ~29 tok/s decode solo at 65K, ~192 tok/s aggregate at 8-way. Rejected AWQ/FP8 weights (72% EOS failure), ExLlamaV3 (3-6x slower), and PP=2 (infeasible on 2 GPUs).

33–48 GBunsloth_Qwen3.8-27B-NVFP4
SHARE YOUR SETUP

Two ways to contribute.

Choose the path that fits you. Both routes go through human review before anything is published.

01 · ANONYMOUSUse the Tally form No account or GitHub sign-in. Best for a one-off contribution.02 · CONNECTEDUse dashboard or an agent Keep a private history, follow review status and update your setup later.
Keep your setup current with an agent.Connect an MCP-compatible assistant to prepare a reviewed update from the hardware and workflows you approve.Discover Agent / MCP