The hardware-first path
Your machine is under‑used. Let's fix that.
Tell us what you're working with. We'll show you every model it can run — and what it would feel like to run them.
Your machine
~2 seconds
or choose a rig
●
All detection stays local in your browser · nothing leaves this page
Your envelope
This is what your machine can hold in memory at once.
Memory decides which models fit. Bandwidth decides how fast they run. The rest is context window and quantization trade‑offs — we'll show you all of it.
—
—
GPU
—
VRAM
—
—
Bandwidth
—
—
Class
—
What fits
Your machine fits models across two categories — chat and code.
Text inference is the only community-verified modality today. Every count below is real: models that fit this exact rig at 4-bit.
Models in Chat that fit your rig.
Sorted by speed. Click any card to see a live simulation of what it would feel like to run.
Test drive
See what it actually feels like.
Pick a model above to see its real pool speed on this rig.
—
— / — GBweights + KV cache
0.0
tokens / sec
live preview
—
Your cohort
People with your rig are running these right now.
Community pool data on this exact rig — aggregated from verified single-stream runs.
Top 5 on your rig
by run count · community pool
Cohort stats
verified pool · snapshot
The bigger picture
This isn't a catalog. It's the operating system for local AI compatibility.
Benchmarks are community-submitted, verified single-stream runs from the public pool.
$
curl -sSf bestmodel.run/sh | sh