A web-based version of llmfit — find what LLM models run on your hardware, right in your browser.
🌐 Live Demo: https://rafaelmaza.github.io/llmfit-web/
This repository was AI-generated as a web port of the original llmfit Rust CLI tool.
All credit for the core algorithm, scoring logic, and conceptual design goes to AlexsJones, the creator of llmfit.
This project reimplements the llmfit scoring engine in JavaScript to make it accessible via a web interface. The original Rust implementation remains the authoritative reference.
llmfit-web brings the model-scoring engine from the original Rust CLI tool to the web, allowing anyone to input their hardware specs and get instant recommendations for which LLM models will run well.
- 50+ GPU models — NVIDIA RTX/A-series, AMD RDNA/MI-series, Intel Arc, Apple Silicon
- 206 LLM models — From 812K to 753B parameters, including MoE architectures
- Multi-dimensional scoring — Quality, Speed, Fit, and Context analysis
- Smart quantization — GGUF and MLX quantization hierarchies with automatic optimization
- Real-time ranking — Instant results, no backend needed
- Mobile responsive — Works on phones, tablets, and desktops
- Select your GPU (or enter custom VRAM)
- Choose your system RAM (8GB to 128GB)
- Pick your use case (Coding, Reasoning, Chat, Multimodal, etc.)
- Get ranked recommendations — Top models scored by fit, speed, and quality
The scoring engine analyzes:
- Memory fit — Does it fit in VRAM, need CPU offload, or use MoE expert offloading?
- Speed — Tokens/sec based on backend (CUDA, ROCm, Metal, CPU) and quantization
- Quality — Parameter count, model family, quantization penalty, task alignment
- Context — Context window availability vs. use-case requirements
- GPU — Model fully loaded into VRAM (fastest)
- MoE Offload — Active experts in VRAM, inactive experts in system RAM
- CPU Offload — Partial GPU, model spills to system RAM
- CPU Only — Entire model in system RAM (slowest)
- Perfect — Recommended memory met on GPU
- Good — Fits with headroom
- Marginal — Tight fit or CPU-only
- Too Tight — Does not fit
llmfit-web/
├── src/
│ ├── scoring.js # Core multi-dimensional scoring engine
│ ├── gpus.js # GPU lookup table (NVIDIA, AMD, Apple, Intel)
│ ├── utils.js # Quantization hierarchy, speed calculations
│ └── models.js # Model database (206 models)
├── test/
│ └── scoring.test.js # Unit + integration tests (30 passing)
├── index.html # Web UI (production-ready)
├── server.js # Local dev server (Node.js)
├── package.json # Dependencies
└── README.md # This file
Start the dev server:
npm startThen visit:
- Local: http://localhost:8000
- Network: http://:8000
Or open index.html directly in a browser (works without a server).
Run the test suite:
npm testCurrent: 30/30 tests passing ✅
- 23 unit tests (scoring engine, memory fit, speed, quality)
- 7 integration tests (real-world hardware + model combinations)
This project is static HTML + JavaScript — no backend, no build step, no runtime dependencies.
Deploy anywhere:
- GitHub Pages (built-in)
- Netlify / Vercel (drag-and-drop)
- Cloudflare Pages
- S3 + CloudFront
- Or just serve the files from any web server
If you want a short GIF for the README, the repo includes a Playwright script that drives a headed Chromium session and records a demo video.
npm installLocal machine with a GUI:
npm run record:demoHeadless Linux server (uses a virtual display):
sudo apt-get install -y xvfb
xvfb-run -a npm run record:demoThe video will be saved under:
artifacts/videos/
This repo includes a script that captures PNG frames via Playwright and encodes an animated GIF using gifenc + pngjs:
npm run generate:gifThis writes demo.gif in the repo root.
206 open-source models included:
| Category | Count | Parameter Range |
|---|---|---|
| Tiny | 31 | <1B |
| Small | 44 | 1-7B |
| Medium | 66 | 7-20B |
| Large | 37 | 20-70B |
| XLarge | 28 | 70B+ |
| MoE | 33 | Mixture of Experts |
Source: Extracted from HuggingFace via the original llmfit project.
Use the scoring engine programmatically:
import { scoreModels, UseCase } from './src/scoring.js';
import { getAllModels } from './src/models.js';
const hardware = {
gpuVramGB: 24,
systemRamGB: 64,
backend: 'CUDA'
};
const useCase = UseCase.Coding;
const results = scoreModels(getAllModels(), hardware, useCase);
console.log(results.slice(0, 5)); // Top 5 modelsOriginal Project: llmfit by AlexsJones
This web port was AI-generated and released under the same license as the original project.
MIT License (matching the original llmfit project)
Issues and PRs welcome! If you find a bug or have a feature request, please open an issue.
For major changes to the scoring algorithm, consider contributing to the original llmfit project first.
