CONSENT CONSOLE/MK-V
DEFAULT

Telemetry consent. Operator-grade.

We capture only the signals we need to keep the site running, understand which content earns reads, and credit referral partners. You decide what stays on. Default is strict opt-in.

Privacy Policy →Terms →
JURISDICTIONOutside regulated jurisdictionsFRAMEWORKNo regional opt-in framework applied

COMPLIANCE FRAMEWORKS RECOGNIZED

GDPREU / EEA
CCPACalifornia
LGPDBrazil
PIPEDACanada
ePrivacyEU Directive
Strategia-X
L
-6dB
C
-1dB
R
-3dB
Technology Trends

A fallback hid a dead feature for months

Rocky ElsalaymehAug 27, 20264 min read768 words
Technology TrendsOP-3855

A fallback hid a dead feature for months

PUB·4 MIN·768 WORDS

In May I told readers that the retrieval stack in Team-X was "faster than I expected." The component I was praising had never executed. An audit this August found that the sqlite-vec path in Team-X, an open-source, local-first desktop app for running AI-agent organizations, threw an error on every call. A catch block absorbed the error and a brute-force fallback answered every query. Nobody noticed, because the answers were right.

That is the executive problem. The failure did not show up as an outage. It showed up as nothing at all, and nothing is the one status no dashboard flags.

What a silent fallback costs you

The changelog for the fix lists five defects in the dead path, four of them independently fatal: a migration missing from the journal, an index collision, an INSERT into a table defined nowhere, and an extension that was never loaded. The fifth was a query that never used the caller's vector. Every call threw no such table: embeddings_vec.

The README meanwhile advertised "sqlite-vec embeddings." The changelog's correction reads, "It never had them." The migration header had promised O(1) search. The code that ran was a linear scan.

What the documents saidWhat the code did
README: sqlite-vec embeddingsEmbeddings stored as BLOBs, ranked by brute-force cosine similarity in-process
Migration header: O(1) similarity searchLinear scan over every stored chunk
Catch block: warn and fall backWarning on every retrieval, so no signal separated healthy from broken

This is a known failure mode. A 2014 USENIX OSDI study of 198 randomly selected, user-reported failures in five distributed systems concluded: "We found the majority of catastrophic failures could easily have been prevented by performing simple testing on error handling code." The error handling is where the risk lives, and it is the code leadership never reviews.

Python's own design guidance, PEP 20, puts it in two lines: "Errors should never pass silently. Unless explicitly silenced." A fallback that does not say which path served the request is silence by default.

You probably do not need the vector database

The second lesson is about spend. Team-X ran on exact in-process search the whole time, and it served every query. The pgvector documentation explains why that is a respectable default: "By default, pgvector performs exact nearest neighbor search, which provides perfect recall." An approximate index "trades some recall for speed."

The FAISS guidance says the same thing from the other side: "The only index that can guarantee exact results is the IndexFlatL2 or IndexFlatIP." Buying an index means accepting inexact answers, plus the cost of building and rebuilding it.

The repo does its own arithmetic. At 10k chunks and 768 dimensions, a full scan is about 7.7M multiply-adds per query, and it grows linearly. A code comment says that below 4,096 vectors a full scan costs "single-digit milliseconds." That is a comment, not a published benchmark. I found no latency figures in the repo, so I will not give you one.

How the decision is made in code

On main, Team-X now has an IVF approximate index. It engages only when a company's corpus reaches 4,096 stored vectors. Below that, retrieval stays exact. Probing every partition reproduces brute force exactly, so the index is a relaxation of the exact path and not a different ranking. That makes the exact path the ground truth for measuring recall.

The changelog reports 93.3% recall at 6 probes and 97.5% at 10 on a synthetic fixture. The test file itself enforces a 90% floor, and I could not find the four published figures in it. I would not plan capacity around either number. The ANN-Benchmarks paper frames the right habit: measure "the performance and quality achieved by nearest neighbor algorithms" on your own data.

What I would ask your team on Monday

  • Which fallbacks exist? List every place a failure is caught and a substitute result is returned.
  • Which path answered? Require a counter on the result, not a log line. A fallback rate of 100% should be impossible to miss.
  • Can the primary path fail a test? Break it on purpose and confirm something goes red.
  • What is the corpus size? Set an explicit threshold before approving any vector infrastructure.
  • Do documents and code agree? Team-X now asserts that its README migration count, the migration files, and the journal all match.

I wrote the full engineering account, with the catch block and the thresholds, in Our vector database never ran. Retrieval was fine. The short version for operators: a system that degrades gracefully will also hide its own decay, so make it say which path it took.

-Rocky

#TeamX #SilentFailures #Retrieval #EngineeringDreams #StrategiaX

Originally published on Team-X Blog.

Team-X Silent Failures Retrieval Vector Search Engineering Leadership Reliability Open Source

/Rocky