Tripwire
Bait inside your app: canary database rows, canary keys, decoy routes and files. When one is touched you hear within seconds, a reversible shield blocks the attacker's path, and an AI explains the break-in and drafts the code fix.
We build AI products for security, private home AI, developer tools, custom agents and automation. Our open research lab measures every idea on real models before it goes anywhere near a product.
Each one solves a single job end to end. The status label on each card is where it really is today.
Bait inside your app: canary database rows, canary keys, decoy routes and files. When one is touched you hear within seconds, a reversible shield blocks the attacker's path, and an AI explains the break-in and drafts the code fix.
A personal AI agent that runs on a small model at home: files, notes, maths and everyday jobs, with every risky command waiting for your yes. Your data stays in the house; heavy work can go to the cloud when you allow it.
Audits a Python codebase for missing and stale docstrings, then writes them and a README for you to review. Made for code nobody remembers writing.
AI agents built around one business's real workflow, starting with recruitment: an assistant reads each new application against the brief, scores it and writes a two-line why or why not. Nothing reaches a candidate or client without a person checking it.
We rebuild dated small-business websites into fast, simple, good-looking ones, and show you a free before and after of your own site first.
For plumbers, electricians and cleaners: every missed call gets an instant text back, quotes get followed up, and happy customers get asked for a review.
Sixteen open experiments on making large language models smaller, faster and cheaper to run, most of them on Qwen3-8B with held-out test text. The negative results are published too.
Qwen3-8B, 22.5% of tokens escalated to bf16: 14.17 ppl vs 15.01 escalating at random (3-bit alone 16.44).
Qwen3-8B at 48k tokens: GPU-only runs out of memory; split 215 ms/token vs 294 CPU attention, 2,201 fetching the cache; same tokens.
Qwen3-8B, 3.5-bit KV cache: 18.8 ppl vs 43–55 for random heads, and beats uniform 4-bit (19.85) with 12% less cache.
Qwen3-8B, all layers 4-bit: 12.26 ppl vs 12.03 FP16 and 16.45 for standard rounding.
Qwen3-8B 3-bit: 14.48 ppl vs 17.47 plain rounding; 4-bit: 12.51 vs 12.92 (bf16 12.03).
Qwen3-8B at 3.5 bits/weight: 13.64 ppl vs 15.38 for random layer choice (3-bit 17.47, 4-bit 12.92).
Qwen3-8B, half the prompt kept: 9.09 ppl vs 9.76 keep-recent vs 8.51 full prompt.
Qwen3-8B: 3.3x faster on code edits, 1.9x on quoting, 1.2x on free text.
Negative result on Qwen3-8B: best 11.49 ppl vs sinks 11.36 at 256 entries (10.10 vs 10.07 at 512). Diagnosis in the repo.
Negative result on Qwen3-8B: at 2.0x smaller SVD is 43.3 ppl, int8 at 1.97x is 11.7 (baseline 11.71).
Qwen3-8B, 200 held-out passages, quarter of the cache kept: 11.33 ppl vs 11.70 for H2O (full cache 9.35).
Qwen3-8B, 512-token GPU budget: 8.89 ppl vs 9.99 sinks; the late oracle choice gets 8.88 (full 8.72).
Qwen3-8B, half the FFN per token: 23.4 ppl vs 33.0 static pruning and 51.0 random towers (dense 12.0).
5 seeds, target 90%: ACI covers 89.9% vs 74.7% for split conformal; the price is wider intervals.
100k points: 58–534 ppm rank error from 119 stored numbers, vs 3,114–50,018 ppm for a random sample of the same size.
Test 2.325 nats/char vs 2.558 for a bigram model (108k parameters, 45 s on a CPU).
Perplexity: lower is better. Each repo has the method, the baseline it was compared against, the full results table and its known limits.
They come from the lab, and they apply to every product we ship.
An idea gets a baseline, held-out test data and a written result before it gets a feature. If it loses, we say so.
Where a small local model can do the job, the data never leaves the machine. Cloud models are the exception, not the start.
Few dependencies, static pages where they'll do, software that runs on modest hardware. Less to break, less to pay for.
Changes ship with tests, risky actions wait for a human yes, and live fixes are built to be reversed.