I've really gotten into local LLMs lately. I'd love more articles on them. I use average models like gemma3:4b or minstral 7b, ones that fit in 8gb VRAM. I feel like they don't work too well with follow up questions sometimes and admittedly get simple stuff blatantly wrong, on a level where even free AI online does magnitudes better. I would love more articles about tweaking the system prompt and helpful tools for making slightly dumber local models work well.