Q4_K_M · 4K context
LFM2.5-VL-3B 3B dense model · estimated Q4_K_M memory about 3GB
It is a roughly 3B vision language model designed to quickly process single, short requests such as OCR, document layout, and image description. Using a separate DSpark draft model speeds up decode on supported devices.
- Estimated Q4_K_M memory
- About 3GB
- total parameters
- 3B
- active parameter
- 3B
- native context
- 32K
Start with a model
Find hardware for this model
Also running document search?
Reserve this for embedding and reranking models sharing the same GPU or unified memory. Leave it at none if they run separately on the CPU.
Matching devices 44· 5 shown
Comfortable fit · Q4_K_M · Headroom 10.6GB
₩1,180,000
46.1 ~ 51.5 tok/s
Comfortable fit · Q4_K_M · Headroom 10.6GB
₩1,525,000
44.7 ~ 50.2 tok/s
Comfortable fit · Q4_K_M · Headroom 10.6GB
₩1,669,000
60.5 ~ 67.4 tok/s
Comfortable fit · Q4_K_M · Headroom 17.6GB
₩1,925,000
46.1 ~ 51.5 tok/s
Comfortable fit · Q4_K_M · Headroom 17.6GB
₩1,950,000
44.7 ~ 50.2 tok/s
Price is the range midpoint, with a basic host added for GPU-only prices. Speed is an estimate for the same model and 4K input.
If you want to choose equipment again
You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.