I released loci v0.6.0 on 10 September. The tokenizer, the file walk, the index build, BM25, the char n-gram matrix and the score fusion all moved behind a compiled Rust extension. The base install is now numpy and nothing else.
A one-shot loci ask against a single scope went from 5.19s to 0.97s.
That is not the number I care about. This one is:
A 39-question benchmark recorded against 0.5.0, capturing every routing verdict, scope set, hit ordering and score, replays 39 out of 39 identical against 0.6.0. The eval harness output is byte for byte the same.
A rewrite that changes the answers is not a port. It is a different tool wearing the old version number, and you find that out from a bad answer six weeks later.
Why touch it at all
loci routes a question to one project, then searches inside that project only. Everything about that design assumes you will ask it things constantly, in the middle of doing something else. Delroy, my agent harness, shells out to loci ask --json mid conversation.
Five seconds per question is a tool you stop reaching for.
Here is where the five seconds went, measured before I wrote a line of Rust:
| 0.5.0 | 0.6.0 | |
|---|---|---|
loci ask, one scope |
5.19s | 0.97s |
| ranker load, largest scope | 935ms | 0.033ms |
| tokenizer, 91k chars | 14.6ms | 4.4ms |
| file walk, 15 scopes | 46.2ms | 20.8ms |
| index build | 14.4s | 9.1s |
Two of those five are not Rust
I want to say this before anyone quotes the first row at me.
The 5.19s included torch loading bge-small. 6.2 seconds of a 7.77 second query, on the run I profiled, was the encoder booting. That is onnxruntime now, loading the ONNX graph BAAI publishes directly. No Rust involved. The deeper reason was not speed anyway: ask fans out across scopes in a thread pool, every thread reached torch, and after three locks closed a diagnosed deadlock, Delroy still measured a segfault. The class of bug was never the check-then-act race. It was native threading, and removing the native threading removed it.
The 935ms was joblib.load on 35MB of pickled scipy, paid on every one-shot invocation that touched that scope. The replacement is a file format rather than a language. rankers/<scope>.lex stores the fitted BM25 and char-gram models with sorted string tables that get binary searched where they lie, and CSR columns renumbered into sorted order so a column id is its own position. Opening it is a page fault instead of a deserialization. I would have got most of that win writing the format in Python.
Rust earned the tokenizer, the walk, the index build and the fusion. It did not earn the headline. If you take one thing from this post and you are staring at your own slow Python, it is that the profiler tells you which of those two situations you are in, and the answer is very often a pickle or a model load rather than the arithmetic everyone wants to rewrite.
What stayed in Python, deliberately
router.py did not move. Neither did any of the fourteen calibrated constants.
The constants cross the boundary as function arguments rather than living in the Rust, because evals/ sweeps them and a sweep must not need a cargo build. search() takes its four fusion weights, its length saturation, its score floor and its three gate parameters as plain floats every call. Recency is precomputed on the Python side too, so $LOCI_NOW still works and no date parsing happens in the extension.
The whole thing is a Python package with one compiled module inside it, not a Rust package with Python bolted on. About 2,000 lines across two crates, loci-core holding the logic and loci-py holding the pyo3 surface.
The fusion releases the GIL for its duration. Holding it would serialise exactly the work the thread pool exists to overlap.
The one thing I could not port
TfidfVectorizer(max_features=60000) keeps the 60,000 most frequent character n-grams. I went to reproduce that and found there was nothing to reproduce.
The cut lands where corpus counts are 1 to 5, and tens of thousands of n-grams share those counts. sklearn breaks the tie with (-tfs[mask]).argsort(), and numpy’s default sort is not stable. Which n-grams survived was decided by sort internals.
On a real corpus: 24,281 of one scope’s 68,418 n-grams tied at a count of 1. Roughly 26% of that scope’s retained vocabulary was arbitrary.
You cannot write a specification against that. Breaking the tie deterministically instead, which is the obvious move, shifted char-gram scores by up to 2.6e-02 and reordered the top 5 for 15 of the 39 benchmark questions. So I removed the cap. The terms it was discarding were singletons, which add columns to the matrix and almost no non-zero entries: measured at +5% matrix bytes, with no change in fit time or query time.
Verified on one fixed index: routing identical, retrieval identical, 45 of 45 evidence chunks returned, the justifying file first for 29 of 44 and in the top 3 for 36 of 44.
That is the only deliberate behaviour change in the release, and it exists because the old behaviour had no definition, not because I found something better.
How I knew nothing moved
Before touching anything, I recorded what 0.5.0 answered.
parity/record.py invokes the CLI as a subprocess rather than importing loci. That is on purpose. The CLI and the MCP server are the only two surfaces the port had to preserve, Delroy never imports loci, and an oracle should grade the surface rather than the internals behind it. A non-zero exit gets recorded rather than raised, because “0.5.0 fails this way” is part of the contract.
parity/check.py diffs the replay under two policies:
- Exact for the decision. Verdict, scope set, ordering, mode, reasons. A routing answer is a choice, and a choice is either the same choice or a regression. Ordering counts, because an agent reads the first hit.
- 1e-6 for the score. Floating point reassociates when the code computing it changes, and a difference in the seventh decimal is noise. A difference big enough to move an ordering shows up under the exact policy instead.
Presentation fields are ignored by suffix, so dropping torch’s progress bar off stderr does not read as a regression. The failure that file exists to prevent is a comparator that passes everything, which would report a clean port that is not one and let every later phase build on the report.
404 tests pass. scikit-learn, rank-bm25 and joblib are still dev dependencies, kept as the reference the parity tests compare against. Deleting the old implementation would delete the proof. Same reason sentence-transformers stays: tests/test_embed.py checks the ONNX encoder against it, and the nine floors in calibration.json were fitted to torch weights.
Which is also why the encoder does not use the model_O1 through model_O4 or the quantised variants sitting in the same repo. An optimised graph drifts from the weights the floors were fitted against, and that is the one thing this module may not do.
What shipping it found
The tradition continues. Five commits landed within hours of the tag, and four of them would have reached a user.
A base install with vectors on disk crashed. The semantic tier is an optional extra, but a data directory that already has embeddings loads them with numpy and then dies importing the provider. ModuleNotFoundError: No module named 'onnxruntime', exit 1, no answer. 0.5.0 had the identical crash one import earlier, on sentence_transformers, and survived only because nobody had run a base install against a store built with the extras. Vectors without an encoder now degrade to lexical search, and loci doctor names the missing package.
Separating the reference tests by pytest marker does not work. importorskip runs at collection, so deselecting by marker still imports scikit-learn into the process. They needed their own file and their own run.
The sdist did not contain the LICENSE that the package metadata had been claiming since 0.1.0.
The walk parity test was comparing the wrong two things. File contents belong against the Rust, ordering belongs against the Python shim, and one test asserting both against one side proves less than it looks like.
Upgrading
A reindex is required. The lexical rankers changed format and one ranking constant is gone, so an older install keeps working but gets slower and answers slightly differently until it rebuilds.
pipx upgrade loci-mem
loci index --force # writes rankers/<scope>.lex; the old .joblib is ignored
loci doctor now says so rather than leaving you to find out. A scope with no usable .lex refits in process on every single question, about 1.5s per scope per invocation, against 0.03ms for a cached one.
INDEX_VERSION stays at 3. Nothing about what a token is changed, and text.rules_signature() returns exactly what it did in 0.5.0. The old rankers/*.joblib files are dead once you reindex and can be deleted.
Wheels ship for macOS, Linux and Windows, abi3 for Python 3.10 and up. Building from the sdist needs a Rust toolchain, 1.75 or newer.
pipx install 'loci-mem[all]'
loci setup
loci ask "why was the session cookie dropped on localhost?"
What has not improved
The limits from the v0.2.0 post are still the limits, and a faster implementation of the same router does not touch any of them.
Every constant was fitted against repositories belonging to one person. Me. The parity work proves 0.6.0 agrees with 0.5.0, which is a statement about a rewrite and says nothing about whether 0.5.0 was right. A question derived from a code body still routes at roughly half the rate of one derived from indexed prose. A large prose-heavy project still over-attracts.
So the ask is unchanged, and the run is now under a second:
loci eval
It asks only questions whose correct answer is known by construction, so it needs no labelling and no setup. Tell me what it says about your projects rather than mine.
Comments
Almost there
Choose how your comment should appear.