NewMeet the LLM compatibility engine

Your machine.
One clear shortlist.

Prism finds the local models your PC can actually run,
with the right quantization, memory cost, and honest limits.

Prism dashboard

Analyze your machine.

Enter your GPU, VRAM, CPU, and RAM. Prism returns a ranked shortlist of models that actually fit — with quantization, memory cost, and honest limits.

Your hardware

Click to pick each field — search if the list is long.

Compare what actually runs.

RTX 4070 · 12 GB · 32 GB

Speed and memory figures are planning estimates, not measured benchmarks.

Compare models →
On 12 GB VRAM
Phi 4 14BMicrosoft
Parameters
7B
14B
70B
Best quant
Q8_0
Q4_K_M
Q4_K_M
Est. VRAM
10.2 GB
10.3 GB
41.2 GB
Est. speed
~23 tok/s
~38 tok/s
~1 tok/s
Context
32K
16K
8K
Verdict
Comfortable
Comfortable
Won't run

Questions, answered.

Not yet. Throughput values are planning estimates derived from model size, quantization, VRAM class, and CPU offload. They are labeled everywhere they appear.

No. No account required. You enter your hardware manually — nothing is installed, uploaded, or granted access to your machine.

The catalog syncs from the public Ollama library (daily when CI is enabled). VRAM and speed figures stay planning estimates — formula-based with published Ollama file sizes when available.

For sizing, yes — enter the unified memory total as both VRAM and RAM. Chip-specific throughput calibration is still in development.

A ranked shortlist of local models with the right quantization, memory cost, runtime, context size, estimated speed, and the main limitation to expect.