Local LLMs for meeting notes: we tested 0.8B to 9B on real meetings

We ran Qwen3.5 0.8B–9B, Gemma 4 E2B and LFM2.5 on real AMI meetings on a 16 GB Mac, and measured memory, speed and how many human-noted decisions each one caught.

TriCode Studio · · 9 min read

local llmllama.cppqwenmeeting notesbenchmark

We make Murmur for Mac, an app that records meetings and writes the notes without anything leaving your computer. We wanted it to write a proper summary too, using a language model on the same machine: no cloud, no API key. Before building that, we needed straight answers to three questions. Which small models fit on the laptops people actually own (8–16 GB)? How long do they take? And how much do they miss compared with a person taking notes?

So we measured. This post covers the set-up, the numbers, a surprise from the biggest model, and the design we shipped because of them.

The short version

  • Speed and memory are not the problem; accuracy is. Every model up to 4B summarised a 39-minute meeting in 32 seconds or less on an M4’s GPU, in under 6 GB.
  • The best small model found about 4 in 10 of the decisions a person noted. Usable, but not something you’d trust without checking.
  • Asking for a citation behind every point made the 4B model consistent on a clearly-run meeting. In one pass it either found 3 of 4 decisions or none. Chunked, with each claim checked against a transcript line, it found 3 of 4 on both runs and used 20% less memory. On a looser, discussion-heavy meeting it still struggled.
  • The 9B did worse than the 4B. It took our “only list what was agreed” rule so literally that it filed every decision as an open question.
  • CPU-only needs more memory than GPU, because llama.cpp keeps a second, repacked copy of the weights.

The set-up

  • Machine: Apple M4 (4 performance + 6 efficiency cores), 16 GB RAM, macOS 26. CPU-only runs used 4 threads, as a stand-in for a Windows laptop without a usable GPU.
  • Runtime: llama.cpp b11371 with llama-server, the way the app runs it: a helper process on localhost. 32K context (room for about a two-hour meeting), flash attention on, thinking turned off.
  • Meetings: two real meetings from the AMI corpus, plus a long case made by joining two: ES2004b (39 min, 7,286 prompt tokens), IS1009b (34 min, 6,258 tokens) and ES2004a + ES2004b (57 min, 10,166 tokens). We transcribed them with Murmur’s own speech recognition, speakers labelled, so the transcripts carry real recognition errors rather than AMI’s clean manual text.
  • Prompt: one fixed prompt for every model. Use only what the transcript says, then write four sections: Summary, Decisions, Action items, Open questions.
  • Sampling: each maker’s recommended settings from the model card, two runs per model (seeds 1 and 2).

One early lesson: our first pass used greedy decoding (temperature 0), and it sent Qwen3.5-0.8B into a 30-line repetition loop. That’s unfair to small models, so we threw those runs away and used the recommended sampling instead.

How we graded

AMI comes with notes written by people for each meeting. We took the four decisions in each meeting’s notes and marked every summary as having caught each one (1), partly (½) or missed it (0). We also counted clear errors: an invented fact, a statement opposite to what was said, or something still undecided listed as decided.

It’s a strict test. The annotators summarised a whole series of meetings, and in ES2004b the decisions are implied by the discussion rather than announced. But the transcripts do contain every decision we graded against: the phrases are there (“eliminate… the teletext”, “the corporate image is being maintained”).

One pass: the whole meeting in one prompt

Model Download Memory needed (GPU / CPU only) 39-min meeting (GPU / CPU only) 57-min meeting (GPU) Decisions caught (of 8, mean of 2 runs)
Qwen3.5-0.8B 0.5 GB 1.3 / 1.5 GB 6 s / 27 s 10 s about 1¼ (16%)
LFM2.5-1.2B 0.7 GB 1.3 / 1.9 GB 7 s / 31 s 12 s about ½ (6%)
Qwen3.5-2B 1.2 GB 2.2 / 2.9 GB 16 s / 51 s 19 s about 3 (38%)
Qwen3.5-4B 2.7 GB 4.4 / 6.4 GB 32 s / 2 min 7 s 43 s about 3 (38%), but each run caught 0 or 3 of 4
Gemma 4 E2B 3.4 GB 5.8 GB / not run 30 s / not run 42 s about 2 (25%)

“Memory needed” is llama.cpp’s own total of weights, context and compute buffers at 32K context, across GPU, host and the CPU copy. What each model did:

  • Qwen3.5-0.8B gets the gist right, and its details are real (we checked each one against the transcript). But it can’t pick out decisions: it repeated the previous meeting’s decisions as this one’s, and one run wrote “None” under every heading. Good for a short overview, nothing more.
  • LFM2.5-1.2B dropped the required sections in two of three runs. Its licence also doesn’t cover commercial use by organisations with $10M or more in yearly revenue, so we couldn’t ship it in an app any company might use.
  • Qwen3.5-2B found about three decisions, with roughly one error per run (for example “rejects video recording”, the opposite of what was said).
  • Qwen3.5-4B was the best when it engaged, and it engaged half the time.
  • Gemma 4 E2B is “2B effective”, but it loads 5.3 GB of weights (per-layer embeddings), more than twice the 2B Qwen’s memory for fewer decisions.

CPU-only costs memory as well as time. Without a GPU, llama.cpp keeps a CPU-optimised copy of the weights alongside the mapped file: the 2B needs 2.2 GB on the GPU and 2.9 GB on the CPU, and runs 3–4× slower.

Chunked, with a source behind every point

A model that finds three decisions on one run and none on the next isn’t something you can ship. So we changed the shape of the task:

  1. Split the transcript into chunks of about 1,500 words.
  2. Ask the model for items from each chunk, each one citing the transcript lines it came from.
  3. Check every citation against the transcript and drop anything uncited.
  4. Merge the chunks’ items into one summary.
Model Memory needed (GPU) Time, 39-min meeting IS1009b decisions (of 4), runs 1 / 2 ES2004b decisions (of 4), runs 1 / 2
Qwen3.5-2B 1.8 GB (was 2.2) 19–22 s — —
Qwen3.5-4B 3.5 GB (was 4.4) 39–65 s 3 / 3 (one pass: 0 / 3) 0 / ½

The 4B became consistent on the clearly-run meeting. It was still weak on the discussion-heavy one, and it sometimes filed a presenter’s feature ideas as action items. Chunking didn’t rescue the 2B: it pulled out junk (“ask whether the plug has come out”) and often skipped the citation. For both, the shorter context cut memory by 17–20%.

The bigger win is for the person reading the notes. Every point now links to the moment it was said, so checking a claim takes one click, and nothing the model can’t trace to the transcript gets through.

The 9B surprise

Qwen3.5-9B looked like the obvious upgrade for 16 GB machines. We ran it on the same meetings with the same chunked pipeline:

Qwen3.5-4B Qwen3.5-9B (Q4_K_M)
Memory needed (8K context) 3.5 GB 6.4 GB
Time, 34–39 min meeting 39–65 s 98–112 s
IS1009b decisions, runs 1 / 2 3 / 3 0 / 0
ES2004b decisions, runs 1 / 2 0 / ½ 0 / 0

It didn’t find a single decision. Our extraction rule says a decision must be something the group agreed, not just suggested, and the 9B applied it to the letter: on IS1009b it produced 17 open questions, every one well cited, and wrote “None” under Decisions. Grounded and useful as a list of questions, but it never told you what was decided, at twice the time and nearly twice the memory. A different prompt might change that; we didn’t tune for it.

What it looks like

Here is the 4B on IS1009b, run inside the shipping app (46 seconds including loading the model), lightly trimmed:

Decisions

  • Suggestions must be specific and not general ideas like “whole house control” ▸ 5:42
  • Teletext feature to be eliminated due to higher management instructions regarding complexity, time to market, and corporate image ▸ 5:42
  • Remote control design should focus on simplicity, intuitiveness, and ease of finding the device ▸ 7:13
  • Target group is basically everybody who has a TV ▸ 33:06

Action items

  • Speaker 2: put slides up ▸ 1:15

Open questions

  • Whether whole house idea is possible within budget ▸ 4:54

Against the people’s notes, the focus, dropping teletext, the corporate image and the target group are all there, about 3½ of 4. It also made mistakes: “put slides up” isn’t an action item worth keeping, and elsewhere it listed a suggested slogan as a decision. Every line links to its timestamp, so these are quick to spot.

What we shipped

  • The default notes stay rule-based. Key points and action items with owners and due dates come from the transcript itself and never invent words.
  • “AI summary (beta)” is optional, chunked and cited, and Murmur picks the model tier from the machine’s memory and GPU:
Tier Model Download Memory needed (GPU / CPU only) What Murmur tells you
Lite Qwen3.5-0.8B 0.5 GB 1.3 / 1.5 GB A short overview. It doesn’t reliably list decisions.
Standard Qwen3.5-2B 1.2 GB 2.2 / 2.9 GB The main points. Misses some decisions; check the sources.
Best Qwen3.5-4B, chunked 2.7 GB 3.5 / about 6 GB About 4 in 10 decisions a person would note. Every point links to the moment it was said.

A rough rule for any machine: total RAM should be at least the model’s need plus about 4 GB for the system and other apps. Without a GPU, use the larger figure and expect it to take 3–4 times as long.

We chose to say plainly what each tier gives up. “AI notes as good as a person” would be easy to write and untrue at this size.

Limits of this test

  • Two meetings, two runs per model, graded by us against AMI’s notes. It’s enough to rank the models and see failure modes, not to give precise accuracy figures.
  • One machine, an M4 with 16 GB. Windows numbers are still to come.
  • English meetings only. AMI meetings are role-played product-design sessions, more structured than many real ones.
  • One prompt, tuned once. A better prompt could move every number here, the 9B’s most of all.

If you’ve had better luck getting small local models to separate decisions from discussion, we’d like to hear what worked.