ProjectsInference
Local fine-tuning and inference
QLoRA fine-tuning of Qwen2.5-Coder-7B on a home RTX 5070 Ti to fix common code-quality issues, plus local inference with Ollama.
- Year
- 2026
- Area
- Inference
- Built with
- Unsloth, QLoRA, Qwen2.5-Coder-7B, Ollama, RTX 5070 Ti
A home-lab project for learning how models really behave on limited hardware: teaching a small open model one narrow skill on a single consumer GPU with 16 GB of memory.
I fine-tuned Qwen2.5-Coder-7B with QLoRA to fix common code-quality issues, and run local inference through Ollama to see where memory, quantization and throughput set the limits. A 12 GB quantized model runs at about 70 tokens per second.