Giving local LLMs a try


Important disclaimer: I will never write prose or personal code using these, I just want to replace Google/StackOverflow, maybe image tagging and reverse search.

The first time §

I used a LLM for the first time a few days ago. Yup, I'm not joking nor exaggerating. I was at work and since we got Gemini for free with our Google account, I finally decided to see what's what.

First, I had a complex question about C++ OOP (Why does a sidecasting dynamic_cast from this base class pointer to this mixin fail? Answer: because I didn't know the public keyword had to be repeated in a multiple inheritance list) and it solved my problem quickly and clearly. I was kinda impressed.

Then today, I asked it to show me a minimal example creating a synthetic 3D image in VTK (for a unit test). Again, no issue and that's using the Flash model that's supposedly far from the heavy hitters Claude and ChatGPT.

But the part that really wowed me is when I asked it to recognize the brand/model of sunglasses the guy in this YouTube video was wearing. It shat itself two times with the YT URL but when I gave it a screenshot, it was quite helpful! Then it gave me useful recs about competing brands, even though I'm not stupid enough to trust Google to not stealthily sell product placement here.

The shopping §

So, having become fed-up with search engines now made of utter fail and AIDS and having just bought myself a Radeon 9070 XT with 16 GB of VRAM (before prices went even madder and without Nvidia's HouseFire™ 12V connector), maybe it was the time to try something?

But here are the rules I laid out:

  1. The software must be simple and easy to sandbox (no access to the network or filesystem).
  2. No "_XxX_UNSLOTH_Q3KXYZ_XxX" model derivative made by Kevin in his mom's basement.
  3. Multimodal, I want to be able to feed it media!

The candidates facing my first constraints were:

  • The de facto standard llama.cpp, performant but supposedly lacking in ergonomics and polish.
  • Sketchy ollama kinda making money on the back of llama.cpp.
  • Finally Pythonware vLLM that's not packaged in the main Portage tree or GURU and doesn't support Vulkan, so no thanks (ROCm is just a massive pain).

A quick win for llama.cpp.

Now, time to pick a model and ho boy what a shitshow it is! The jargon you must understand, the lack of clear data about the accuracy cost of quantization/reduced model size or simply the final memory requirements and the spotty support for shiny new features like TurboQuant, N-gram embeddings or multi-token predictionMTP.

The only two models that retained my attention were the much talked about Qwen 3.8 by Alibaba and Google's Gemma 4, but to fit their respective 27B/26B models in my small GPU, I had to reach for aggressive weight quantization. My understanding is that with quantization aware trainingQAT, Q4 is the floor before output accuracy starts to plummet. Sadly, Qwen isn't offered by Alibaba in Q4, so Gemma it is!

The home setup §

After installing llama.cpp from GURU and looking for a long time at its repo, I decided to never give this poorly developed piece of C++ access to my system. So, I manually downloaded the model files and immediately threw it into my homegrown bubblewrap jail before starting to experiment.

First reactions: that many options and no manpage nor zsh completion!? Where's the fucking doc? What's the difference between llama-cli and llama-completion? Huh, why isn't the --image option working? Why does it output Markdown soup instead of using normal terminal escape sequences? Can't I hide the reasoning part and just get a dumb spinner? Why NIH code a bad line editor instead of using libedit? WHY DOES IT STOP MID-REPLY!?

...I'm starting to understand why people look at ollama.

An hour of pain and a rlwrap wrapper later, I have something that kinda works. Still bad enough that I'm considering an Emacs client instead.

#!/bin/sh
# Issues:
#   Sometimes stops mid-output/thinking, maybe due to implicit --fit on ?
#   llama-cli --image not working, must input /image manually
#   No way to output ANSI bold instead of Markdown ?
#
# Try server with UNIX socket and https://github.com/karthink/gptel
set -eu

cd "$(dirname "$0")"
exec ezbwrap -m llama-cli -n llama-cli -r /usr/share/terminfo \
    rlwrap \
        --ansi-colour-aware \
    llama-cli \
        --simple-io \
        --model   gemma-4-26B_q4_0-it.gguf \
        --mmproj  gemma-4-26B-it-mmproj.gguf \
        --fit     off \
        --prompt 'Output plaintext instead of Markdown and try to provide at
                  least one URL to support your claims' \
        "$@"

Conclusion §

screenshot.png

It's alive and somewhat useful. Nowhere near what hosted LLMs offer and through an unstable software ecosystem (the server might be less janky, though) but I might slowly replace Google with it.

PS: still amazed (and disappointed) we haven't got an Emacs Flymake backend to replace Grammarly by now!