ComingUp ComingUp
S

slopsome.com

Jun 14, 2026 AI & Machine Learning
gpu llm local ai quantization vram

About

Got tired of guessing whether a model fits my card, so I made slopsome.com — a VRAM fit-calculator + real tokens/sec for local and API models. Pick a model, your GPU, and quant, and it tells you if it fits / needs offload / multi-GPU / won't fit, plus rough speed. Just added EXL3 (bits-per-weight) and an advanced panel for KV-cache quant + draft tokens and concurrency, since that's where VRAM quietly blows up. submitted by /u/Defiant_Rough5325 [link] [comments]

Comments (1)

Colin Price Colin Price 1 month ago

how do you plan to monetize this long term