In May I told readers that the retrieval stack in Team-X was "faster than I expected." The component I was praising had never executed. An audit this August found that the sqlite-vec path in Team-X, an open-source, local-first desktop app for running AI-agent organizations, threw an error on every call. A catch block absorbed the error and a brute-force fallback answered every query. Nobody noticed, because the answers were right.
That is the executive problem. The failure did not show up as an outage. It showed up as nothing at all, and nothing is the one status no dashboard flags.
What a silent fallback costs you
The changelog for the fix lists five defects in the dead path, four of them independently fatal: a migration missing from the journal, an index collision, an INSERT into a table defined nowhere, and an extension that was never loaded. The fifth was a query that never used the caller's vector. Every call threw no such table: embeddings_vec.
The README meanwhile advertised "sqlite-vec embeddings." The changelog's correction reads, "It never had them." The migration header had promised O(1) search. The code that ran was a linear scan.
| What the documents said | What the code did |
|---|---|
| README: sqlite-vec embeddings | Embeddings stored as BLOBs, ranked by brute-force cosine similarity in-process |
| Migration header: O(1) similarity search | Linear scan over every stored chunk |
| Catch block: warn and fall back | Warning on every retrieval, so no signal separated healthy from broken |
This is a known failure mode. A 2014 USENIX OSDI study of 198 randomly selected, user-reported failures in five distributed systems concluded: "We found the majority of catastrophic failures could easily have been prevented by performing simple testing on error handling code." The error handling is where the risk lives, and it is the code leadership never reviews.
Python's own design guidance, PEP 20, puts it in two lines: "Errors should never pass silently. Unless explicitly silenced." A fallback that does not say which path served the request is silence by default.
You probably do not need the vector database
The second lesson is about spend. Team-X ran on exact in-process search the whole time, and it served every query. The pgvector documentation explains why that is a respectable default: "By default, pgvector performs exact nearest neighbor search, which provides perfect recall." An approximate index "trades some recall for speed."
The FAISS guidance says the same thing from the other side: "The only index that can guarantee exact results is the IndexFlatL2 or IndexFlatIP." Buying an index means accepting inexact answers, plus the cost of building and rebuilding it.
The repo does its own arithmetic. At 10k chunks and 768 dimensions, a full scan is about 7.7M multiply-adds per query, and it grows linearly. A code comment says that below 4,096 vectors a full scan costs "single-digit milliseconds." That is a comment, not a published benchmark. I found no latency figures in the repo, so I will not give you one.
How the decision is made in code
On main, Team-X now has an IVF approximate index. It engages only when a company's corpus reaches 4,096 stored vectors. Below that, retrieval stays exact. Probing every partition reproduces brute force exactly, so the index is a relaxation of the exact path and not a different ranking. That makes the exact path the ground truth for measuring recall.
The changelog reports 93.3% recall at 6 probes and 97.5% at 10 on a synthetic fixture. The test file itself enforces a 90% floor, and I could not find the four published figures in it. I would not plan capacity around either number. The ANN-Benchmarks paper frames the right habit: measure "the performance and quality achieved by nearest neighbor algorithms" on your own data.
What I would ask your team on Monday
- Which fallbacks exist? List every place a failure is caught and a substitute result is returned.
- Which path answered? Require a counter on the result, not a log line. A fallback rate of 100% should be impossible to miss.
- Can the primary path fail a test? Break it on purpose and confirm something goes red.
- What is the corpus size? Set an explicit threshold before approving any vector infrastructure.
- Do documents and code agree? Team-X now asserts that its README migration count, the migration files, and the journal all match.
I wrote the full engineering account, with the catch block and the thresholds, in Our vector database never ran. Retrieval was fine. The short version for operators: a system that degrades gracefully will also hide its own decay, so make it say which path it took.
-Rocky
#TeamX #SilentFailures #Retrieval #EngineeringDreams #StrategiaX
Originally published on Team-X Blog.
