kvdd.eu · writing

Writing

Notes on evaluating AI and computer-vision systems — how the claims are made, how they fail, and what is worth asking before you commit. Client work is confidential, so nothing here is about anyone's system in particular; the patterns are general.

15 July 2026 · generalisation · ~9 min

Why AI models fall apart on a second site's data

A model scores 0.96 in the demo and lands at 0.75 on your data. Usually the model didn't lie — the evaluation did. What site-level shift actually is, and the six things to ask for.

More to follow.