Applied AI studio · UK

Applied AI, measured and shipped.

We build AI products for security, private home AI, developer tools, custom agents and automation. Our open research lab measures every idea on real models before it goes anywhere near a product.

6products in build or for hire
16public experiments, every number measured
8Bparameter model we test on, on our own hardware
01 / Products

One studio, six products.

Each one solves a single job end to end. The status label on each card is where it really is today.

P01In development

Tripwire

Security for small teams' apps

Bait inside your app: canary database rows, canary keys, decoy routes and files. When one is touched you hear within seconds, a reversible shield blocks the attacker's path, and an AI explains the break-in and drafts the code fix.

Built for Supabase, Firebase, Next.js and Vercel apps.
P02In development

Hearth

Private AI for the home

A personal AI agent that runs on a small model at home: files, notes, maths and everyday jobs, with every risky command waiting for your yes. Your data stays in the house; heavy work can go to the cloud when you allow it.

Completes 55 of 60 runs of our 20-task test suite on a local 8B model.
P03Early access

Legacy Doc-AI

Documentation for old code

Audits a Python codebase for missing and stale docstrings, then writes them and a README for you to review. Made for code nobody remembers writing.

legacy-doc-ai.pages.dev · free audit of a public repo
P04Taking projects

Apex Agents

Custom AI agent systems

AI agents built around one business's real workflow, starting with recruitment: an assistant reads each new application against the brief, scores it and writes a two-line why or why not. Nothing reaches a candidate or client without a person checking it.

Free trial on one live role.
P05Taking projects

Web Studio

Website rebuilds

We rebuild dated small-business websites into fast, simple, good-looking ones, and show you a free before and after of your own site first.

From £300. Free mockup before you decide.
P06Free trial

Trades Desk

Automation for local trades

For plumbers, electricians and cleaners: every missed call gets an instant text back, quotes get followed up, and happy customers get asked for a review.

First three firms get a free trial, then £250 setup and £49 a month.
02 / Research lab

Every claim has a number. Every number has a repo.

Sixteen open experiments on making large language models smaller, faster and cheaper to run, most of them on Qwen3-8B with held-out test text. The negative results are published too.

R01
Quantization · Offloading

Spend precision only where it's needed

Qwen3-8B, 22.5% of tokens escalated to bf16: 14.17 ppl vs 15.01 escalating at random (3-bit alone 16.44).

R02
Inference · Memory

Long context on consumer hardware

Qwen3-8B at 48k tokens: GPU-only runs out of memory; split 215 ms/token vs 294 CPU attention, 2,201 fetching the cache; same tokens.

R03
KV cache · LLMs

Give each attention head its own bits

Qwen3-8B, 3.5-bit KV cache: 18.8 ppl vs 43–55 for random heads, and beats uniform 4-bit (19.85) with 12% less cache.

R04
Quantization · Optimisation

4-bit weights that barely notice

Qwen3-8B, all layers 4-bit: 12.26 ppl vs 12.03 FP16 and 16.45 for standard rounding.

R05
Quantization · Calibration

Protect the busy channels

Qwen3-8B 3-bit: 14.48 ppl vs 17.47 plain rounding; 4-bit: 12.51 vs 12.92 (bf16 12.03).

R06
Quantization · Compression

Spend the extra bit wisely

Qwen3-8B at 3.5 bits/weight: 13.64 ppl vs 15.38 for random layer choice (3-bit 17.47, 4-bit 12.92).

R07
Information theory · Prompts

Shorter prompts, same meaning

Qwen3-8B, half the prompt kept: 9.09 ppl vs 9.76 keep-recent vs 8.51 full prompt.

R08
Inference speed

Speculative decoding

Qwen3-8B: 3.3x faster on code edits, 1.9x on quoting, 1.2x on free text.

R09
Attention · Long context

Do summaries of forgotten tokens help? No.

Negative result on Qwen3-8B: best 11.49 ppl vs sinks 11.36 at 256 entries (10.10 vs 10.07 at 512). Diagnosis in the repo.

R10
KV cache · Compression

Is the KV cache low-rank? Not usefully.

Negative result on Qwen3-8B: at 2.0x smaller SVD is 43.3 ppl, int8 at 1.97x is 11.7 (baseline 11.71).

R11
KV cache · Eviction

Which tokens to keep

Qwen3-8B, 200 held-out passages, quarter of the cache kept: 11.33 ppl vs 11.70 for H2O (full cache 9.35).

R12
Systems · Transformers

Fetch the right pages a layer early

Qwen3-8B, 512-token GPU budget: 8.89 ppl vs 9.99 sinks; the late oracle choice gets 8.88 (full 8.72).

R13
Architecture · MoE

Turn a dense model into experts

Qwen3-8B, half the FFN per token: 23.4 ppl vs 33.0 static pruning and 51.0 random towers (dense 12.0).

R14
Statistics · Reliability

Error bars that survive drift

5 seeds, target 90%: ACI covers 89.9% vs 74.7% for split conformal; the price is wider intervals.

R15
Algorithms · Streaming

Percentiles of a huge stream in 119 numbers

100k points: 58–534 ppm rank error from 119 stored numbers, vs 3,114–50,018 ppm for a random sample of the same size.

R16
Deep learning · Fundamentals

A GPT and its autodiff, from scratch

Test 2.325 nats/char vs 2.558 for a bigram model (108k parameters, 45 s on a CPU).

Perplexity: lower is better. Each repo has the method, the baseline it was compared against, the full results table and its known limits.

03 / How we build

Four rules we don't bend.

They come from the lab, and they apply to every product we ship.

Measure first

An idea gets a baseline, held-out test data and a written result before it gets a feature. If it loses, we say so.

Private by default

Where a small local model can do the job, the data never leaves the machine. Cloud models are the exception, not the start.

Small and fast

Few dependencies, static pages where they'll do, software that runs on modest hardware. Less to break, less to pay for.

Nothing unchecked goes live

Changes ship with tests, risky actions wait for a human yes, and live fixes are built to be reversed.

04 / Contact

Have a job for software? Tell us.

[email protected]