I build AI systems to solve real problems — and learn from where they break.

I’m a technical founder and product-minded engineer. I build agents, evaluation systems, context infrastructure and customer-facing software. My work starts with a real problem, moves quickly into a working product, and improves through close observation of where models, tools and workflows fail in practice.

What I built, what I learned

Each project started with something useful to build for a real user or customer. The most useful lessons emerged when the first version met reality.

Working private system

Crispin

  • AI agents
  • Evaluation systems
  • Notion
Built
A Telegram assistant for organising work in Notion and Google Calendar, with focused agents and 300+ evaluations.
Reality taught me
A model can sound confident without completing the work. I learned to verify tool calls and state changes, then move guarantees into code.
How I built an agent I could trust
Working product

Eieye

  • Next.js
  • PostgreSQL
  • pgvector
Built
A full-stack AI intelligence product that collects HN stories, articles and comment trees, then turns them into a feed, weekly reports and cited research answers.
Reality taught me
Retrieval is only the start. Useful research depends on preserving source structure, inspecting both articles and discussions, bounding the agent, and checking every citation against evidence it actually read.
Making the HN hive mind queryable
v0.1 release candidate

mdReview

  • Go
  • Preact
  • TypeScript
Built
A browser interface for reading Markdown, commenting on documents or selected text, and saving feedback beside the original file.
Reality taught me
Comments can stay with a Markdown file without changing the document or moving it into another service.
Reviewing agent-made Markdown
Published experiment

Chat G&T

  • LoRA fine-tuning
  • Held-out evaluation
  • Structured outputs
Built
An interactive experiment comparing two ways of teaching a small AI model to turn questions into cocktail recipes.
Reality taught me
Fine-tuning learned the response contract more reliably than it learned judgment. It improved structure and removed most recurring prompt context, but did not establish that the answers were better.
Prompting vs fine-tuning in practice
Working pilot

Latently

  • TypeScript
  • Next.js
  • Hono
Built
A tool that gives AI assistants the right project context from Drive, Slack, Notion and files.
Reality taught me
Context is not simply a retrieval problem. Freshness, permissions, source truth and inspectability matter before embeddings or semantic search become useful.
Building reliable context for AI tools

Other things shipped along the way.

01

Gather

Furniture provenance workflow

02

Origin Thread

Physical and digital provenance

03

MushroomRise

AI content automation

04

Lock In Guru

Habit formation application

05

Fliq

Educational discovery application

Customer context is part of the engineering work.

My route into engineering was unconventional. I trained at the Royal Academy of Music, then worked in sales, growth and early-stage companies before learning software through founder problems that needed solving.

That background is useful in forward-deployed work: I listen closely, communicate clearly, work comfortably without a complete specification, and stay focused on whether the system solves the user’s actual problem.

More about my background