I'm a solo engineer who builds systems that improve themselves — multi-agent platforms, evolutionary
code generators, data-intelligence graphs, and generative-3D pipelines. The through-line is the same
thesis this site is named for: technology is art. The work below splits into two worlds —
a Passion side built for the love of the craft, with Mochi OS at its hub, and a
Profession side shipped as professional work, anchored by Nexus.
The system
Everything I build sorts into two worlds: the Passion side — the Mochi OS ecosystem, built for the love of it — and the Profession side — Nexus and the platforms shipped as professional work. Pick a node to follow its connections, or open its case study.
Passion
Built for the love of the craft — Mochi OS at the hub, every tool wired back to it.
Most of what I build runs on GPUs, CLIs, or remote boxes. The ones with a live interface are captured here — running, not mocked. Open any frame for the full case study.
This site's articles are authored by Mochi, an AI persona I built — one of the projects above.
Mochi OS
Architect & sole engineer · 2025–present
A self-improving operating system for software work
Captured from the running system
A multi-agent platform where AI agents build, review, and ship code under enforced quality gates — a daemon serving a git-worktree deploy with guardian auto-promotion, a one-way baseline-ratchet gate system, isolated parallel-agent commits, and a code-graph integrity engine. A living System Map renders the whole OS as a circuit board over real telemetry. The meta-system that builds the others.
System map
Substrate
Lifecycle Kernel
Plugin Host
Deploy Substrate
Frontend & HUD
Senses
Metrics
Overview
Host Telemetry
System Map
Hands
Swarm Engine
Swarm Orchestration
Agent Runtime
Memory Ledger
Will
Control Plane
Living Roadmap
Plan Lifecycle
Operator CLI
Immune
Sentinels & Self-Heal
Guardian Brain
Process Guardian
Reach
Git Sync
Git Projects
Fleet Deploy
Operator Comms
Daemon serves a git-worktree deploy; a guardian auto-promotes master
One-way baseline-ratchet gates block new findings, grandfather old ones
9,000+ unit tests greenMulti-stage promotion gatesLocal-model swarm
Stereopolis
Engineer · 2026
A single photo or video → an explorable 3D Gaussian-splat world
A six-stage pipeline that turns one photograph into an explorable 3D world: panorama → camera trajectory → render → stereo reconstruction → gaussian-splat data → 3D Gaussian Splatting, producing multi-million-gaussian point clouds and TSDF meshes. Validated end-to-end on A100 GPUs from a real photo. It also reconstructs photorealistic Gaussian-splat scenes from real multi-view site captures — a feed-forward reconstruction path plus a COLMAP→gsplat photogrammetry path — and serves them as map-to-splat dives a host map cross-fades into.
Full 6-stage single-image pipeline runs single-GPU on a shared A100
Produced a 1.4M-gaussian .ply + TSDF mesh from a real photograph
Real-capture scan path: a two-plane GPU authoring / GPU-free delivery system with an embeddable tour viewer, plus a faithful COLMAP→gsplat photogrammetry route
The flagship of my professional work: an AI-native B2B workplace for a commerce ecosystem — sourcing, supplier/buyer matching, RFQs, and market intelligence — where trust is the moat. Every answer is sourced, dated, and verifiable, role-adapted (Buyer/Seller/Partner) and bilingual (EN/VI). It is fully sovereign and self-hosted: the agent kernel, knowledge platform, graph-RAG service, Trust Layer, and WebUI are all hand-built — no orchestration framework, no external SaaS on the hot path. Three colocated GGUF models (a 27B chat model plus 4B embedding and reranker) run on a single 48 GB GPU behind one OpenAI-compatible LLM Gateway.
Built every framework primitive in-house — agent kernel, knowledge platform, graph-RAG, and Trust Layer — with zero orchestration-framework dependencies
A Trust Layer that makes every answer sourced, dated, and verifiable, so the workplace can drive real commercial decisions — provenance over product
Runs three colocated GGUF models on a single 48 GB GPU behind one OpenAI-compatible LLM Gateway, with OpenRouter as the only opt-in external fallback
Organised as seven capability planes × eleven product domains over five deployable processes, on Postgres + pgvector and Redis as the only infrastructure
A trade-intelligence system that ingests global bill-of-lading and customs data into a Postgres property-graph, resolves company entities across sources, and derives market signals — supplier risk, under-invoicing detection, alternative-supplier discovery — over a ~270k-shipment / ~60k-company graph. Includes a Company Intelligence Store and source connectors.
System map
Acquisition
Source Connectors
Parallaxis Proxy
Graph Crawlers
Scrape Jobs
Raw Store
Schema & DDL
Raw Stores
Raw Tables
Reference Data
Entity Resolution
Entity Resolver
Auto-Resolve Loop
Entity Dedup
Review Queue
Enrichment Fabric
DAG Framework
Shipment Enrichers
Company Enrichers
Enrich-All Orchestrator
Graph & Communities
Graph Loader
Graph Queries
Communities (LPA)
Graph API
Serving & Presentation
Pipeline Orchestration
Merge Engine
Pipeline API
Frontend (Lumen)
Postgres property-graph over hundreds of thousands of shipments
Cross-source entity resolution + a Company Intelligence Store
A market-signals layer (supplier risk, under-invoicing, alt-supplier discovery)
PythonPostgreSQLGraph algorithmsEntity resolutionWeb data collection
A self-hosted broker that manages a pool of network egress targets with leasing, health-scoring, and multi-armed-bandit routing, feeding the trade-intelligence crawlers. Hardened over multiple audit rounds for lease hygiene, back-pressure, and security.
Lease-based target allocation with health-scoring + cooldowns
TypeScript/NodePostgreSQLMulti-armed banditsDistributed systems
Lease TTL tuningBandit-routedAudit-hardened
swarm
Engineer · 2026
Evolutionary self-improving code generation
Best-of-N candidate generation with in-context parent learning, a multi-persona review panel, bandit model-routing, and a human-gated self-modification scaffold (DGM / ShinkaEvolve-inspired). Runs entirely on local models — sovereignty preserved.
Diverse best-of-N candidates with early-stop + novelty rejection
A deterministic-consensus multi-persona review panel
A human-gated self-modification scaffold (never auto-promotes)
TypeScriptLocal LLMsEvolutionary algorithms
Local-only modelsCandidate archiveBandit-routed
keelloom
Engineer · 2026
Code-graph & index-integrity engine
The engine Mochi OS edits code through: a deterministic code-graph, reference-fact sidecar, incremental authoritative-diff sync, and isolated parallel-agent commits (cowork). A standalone, deployable multi-repo product rebuilt from VibeKeeper — the PM engine inside arobid-nexus (Profession side). Both serve the same mission: governing how AI agents collaborate on code. VibeKeeper runs embedded inside a production B2B platform, enforcing trade-intelligence workflow integrity. Keelloom runs standalone across the Mochi fleet, governing 12+ repos with a same-machine learning loop. Two implementations, one insight: the method had to survive being extracted from its host to prove it was a product, not a feature.
System map
Code-Graph Engine
Indexing & Maintenance
Query & Navigation
Architecture & Insight
Method & Governance
Method Graph & Gates
Cost Meter & Telemetry
Surface & Integration
MCP & Claude Code
Hub, Fleet & Coordination
Hub & Releases
Repo Adoption
Fleet & Learning
Cowork Coordinator
Deterministic code-graph with reference-fact sidecar + scoped-write incremental sync
Isolated parallel-agent commits (cowork) with index integrity
Governs 12+ repos with 35 pre-lock gates, fleet federation, and a same-machine learning loop
A multi-model llama.cpp inference manager on one GPU
Captured from the running system
A FastAPI proxy that fronts llama.cpp, spawning and supervising llama-server subprocesses to serve multiple GGUF-quantised models on a single GPU behind an OpenAI-compatible API (chat, embeddings, rerank). It does VRAM-budgeted LRU load and eviction, tiered model selection, and crash-safe process supervision, with a deep admission layer — tier queueing, per-model circuit breakers, rate limits, quotas, and a context-window guard — plus a built-in dashboard and Prometheus metrics.
System map
API Surface
OpenAI Proxy
Chat
Embeddings
Rerank
Admission
Tier Queue
Circuit Breakers
Rate Limits & Quotas
Context-Window Guard
Registry & Supervisor
Model Registry
Tiered Selection
Spawn & Load
VRAM LRU Eviction
Inference
llama.cpp Servers
GPU / VRAM
Watchdog & Shutdown
Observability
Dashboard
Prometheus
OpenTelemetry
Runs multiple GGUF models on one GPU with VRAM-budgeted LRU eviction, concurrent-load dedup, and an OpenAI-compatible chat/embeddings/rerank surface
Built a layered admission stack — tier queueing, circuit breakers, rate limits, quotas, and a context-window guard that rejects over-n_ctx prompts before they crash llama.cpp
Hardened process supervision and graceful shutdown with cross-platform orphan-prevention (JobObject / pdeathsig / watchdog)
~55k lines of Python467 test files, 40 ADRs13 admission metrics
Virtual Agency
Architect & sole engineer · 2025–present
An AI office-sim where agents autonomously ship real output
An AI-powered office simulation: you play CEO, hiring and managing a team of LLM agents that autonomously collaborate, conflict over decisions, and deliver real runnable artifacts (HTML/JS games, stories, generative art, research). Agents form vector-embedded memories, distil rules, and build relationships under a full economy of salaries, budgets, goals, and routines — running 100% locally via llama.cpp on an RTX 5090 with zero external API calls.
Built a self-running agent engine (perceive → think → act) with vector-embedded memory, distilled rules, and emergent LLM-driven conflict resolution
Shipped a full agency-operations layer: hiring pipeline, salary/budget economy, CEO approvals, and scheduled + emergent routines
Wired the game into Mochi OS as a WSL-native systemd satellite with native HUD monitoring tabs
~38.8k lines across 229 files28 Postgres tables, ~60 endpoints32 project types, 100% local
Fossilize
Architect & sole engineer · 2026
A browser-driven archiver for fully offline website copies
A website-archival tool that drives a stealth Firefox browser to capture modern SPAs into fully offline-functional archives, with parallel asset downloading, SHA-256 verification, pause/resume, and incremental updates. A FastAPI backend with a React dashboard streams live capture progress over WebSocket, and a multi-engine resilience layer grades capture quality (HTML + screenshot-entropy) and auto-recovers failed captures. Runs as a manifest-driven Mochi OS satellite with systemd health probes.
Captures bot-protected SPAs (React/Vue/Angular) into pixel-faithful offline archives with auto-scroll lazy-loading and computed-style extraction
Built a multi-engine resilience layer with deterministic per-domain fingerprints, telemetry-ranked engine selection, and a quality oracle that fails visually-blank captures
Wired as a zero-TypeScript Mochi OS satellite, discoverable and controllable from the OS dashboard
A single-user, fully local pixel-art sprite generator built on the FLUX.2 diffusion model with native GGUF inference, so generation runs entirely on-device with no cloud calls. It owns the whole prompt-to-engine pipeline — text prompt to spritesheet through a seven-step flow (downscale, dither, background removal, grid split, animation) — exporting game-ready assets for Unity, Godot, Aseprite, and Phaser.
Built a native GGUF inference path (no llama.cpp dependency) running a quantised text encoder dequantised one layer at a time to fit VRAM
Shipped a server-authoritative sprite-sheet platform with a job state machine, SSE streaming, cancel/retry/finalize, and 8 engine-native export formats
Implemented a custom sampling engine with 64 samplers and 18 schedules plus a Sampler Lab for XY-matrix comparison
A satellite that turns a single product photograph into a game- and web-ready PBR-textured .glb, built on the TRELLIS.2 structured-3D-latent model with native GGUF inference. A GPU-free control plane (a FastAPI + SSE job queue, SQLite jobs, and a <model-viewer> gallery) runs the GPU engine as an isolated child process for crash isolation, taking the image through sparse-structure generation, structured-latent decode, mesh extraction, quadric decimation, and a 2048² PBR texture bake. Brought up and verified end-to-end on an RTX 5090 (Blackwell / sm_120).
System map
Control Plane
CLI
Job API · SSE
Job Store
Model Gallery
Generation Engine
GGUF Weights
Sparse Structure
Structured Latent
Mesh Decode
Mesh Finishing
DC Remesh
Quadric Decimate
PBR Texture Bake
GLB Export
GPU Substrate
torch · cu130
Triton · sm_120
RTX 5090 / VRAM
Built a GPU-free control plane that drives the GPU engine as an isolated child process — durable SQLite jobs, SSE progress, and a <model-viewer> gallery
Ran TRELLIS.2 with native GGUF inference (Q8_0 + 1536 cascade) to bake 2048² PBR-textured .glb meshes, decimated to a per-preset face budget
Brought the pipeline up on Blackwell: patched Triton sm_120 codegen, prebuilt cu128 wheels, and tamed DC-remesh over-tessellation in mesh finishing
Vietnamese content-moderation ML for a B2B marketplace
A Vietnamese-language content-moderation system for a B2B marketplace, detecting prohibited and NSFW text to meet Decree 147/2024 and E-Commerce Law 2025 compliance. A trained mmBERT-base neural classifier reaches 0.960 P0-F1 on held-out test data, fed by a separate hardened, resumable crawler that runs every sample through fail-closed CSAM/PII drop-filters with no raw-content retention.
Trained an mmBERT-base classifier hitting 0.960 P0-F1 (96.4% recall) and 0.846 action macro-F1 on a held-out test set
Built a crash-safe, proxy-aware crawler with an atomic SQLite ledger, robots-awareness, and a 0.5 RPS cap, live-verified end-to-end
Engineered a fail-closed content-safety pipeline: CSAM/PII drop-filters, HMAC author hashing, and derived-only storage
0.960 P0-F1, 0.846 macro-F1~490K-row unified VN corpus0.5 RPS, zero raw retention
VibeKeeper
Engineer · 2026
AI-agent workflow governance embedded in a B2B platform
The vibe-pm engine inside arobid-nexus: an embedded governance layer that enforces how AI agents collaborate on trade-intelligence workflows. VibeKeeper governs the professional side — ensuring sourcing pipelines, supplier enrichment, and market-intelligence queries follow the platform's trust and compliance rules. It is the progenitor of keelloom (Passion side): the same mission of AI-agent workflow integrity, proven first inside a production B2B system, then extracted and rebuilt as a standalone multi-repo product. The comparison is deliberate: VibeKeeper proves the idea works under production constraints; keelloom proves it generalises beyond a single host.
Embedded governance for AI-agent trade-intelligence workflows inside a production B2B platform
The progenitor of keelloom — same mission, proven in production before extraction to standalone
Enforces trust, compliance, and workflow integrity across sourcing, enrichment, and intelligence pipelines
Large-scale forum data-collection with safety monitoring
A resilient forum crawler with built-in safety/abuse monitoring and a live observability dashboard (frontier, ETA, throughput sparklines, event feed). Hardened against convergence bugs with a durable, fail-open monitoring UI.
A from-scratch CMS built to replace three legacy production sites, running on Node built-ins only — node:sqlite, node:http, node:crypto — with zero runtime npm dependencies in the engine and capability gaps closed by OS services (Caddy, Valkey, libvips) instead of packages. One renderer, used two ways: the same pure (route, data) => Html function serves each request and pre-warms a static mirror. Now live in production on oldbooks.me, veritart.me, and seedy.farm, with an admin console, SSO gating, and a public MCP gateway for AI-authored content.
Zero runtime npm dependencies in the engine — capability gaps closed by Node built-ins and OS services, not packages
One renderer used two ways: the same pure (route, data) => Html function serves requests and pre-warms a static mirror
Replaced all three legacy sites, now live in production with an admin console, SSO, and a public MCP content gateway
node:sqliteTypeScriptCaddy
Spoilers of the Three Kingdoms
Engineer · 2026
A local-LLM dynastic narrative game
An LLM-driven dynastic narrative game set in 184 AD, opening on the Han empire's far southern frontier in Cửu Chân — built on a purpose-written game engine (造化 / Zaohua) in Rust. A Rust orchestrator runs an eight-stage dialogue pipeline over a local Qwen3.8-27B model via llama.cpp — no cloud calls — streaming every NPC conversation live. A SolidJS + Phaser client renders the focused-stage shell: a 229-city campaign map, real-time-with-pause battles, dialogue, and a full economy simulation. Backed by a curated primary-source research corpus spanning bilingual (Vietnamese + English) historical analyses of Han-era governance, military systems, and regional cultures — each research document fact-checked against period sources and cross-referenced with modern scholarship.
Rust orchestrator streams an 8-stage dialogue pipeline over SSE against a local Qwen3.8-27B model, no cloud calls — the 造化 (Zaohua) engine
A 229-city campaign map with turn-based strategy, real-time-with-pause battles, and a living world economy
18 of 23 client app slots live in a SolidJS + Phaser focused-stage shell, driven by a registry-derived test suite
A bilingual primary-source research corpus (Vietnamese + English) with 90+ documents spanning governance, military, economy, and regional cultures of the Han era
RustSolidJSPhaserLocal LLMPrimary-source research
Mochi Auth
Engineer · 2026
First-party single sign-on for the Mochi realm
First-party, branded single sign-on for the Mochi realm — log in once at the portal and reach every gated app with no per-app re-login. Sits behind Caddy as a forward_auth checkpoint, with password + TOTP factors and host-only cookies for SSO without a shared-domain cookie. Live in production, with a full operator console: a system launcher with live health, sign-in metrics, security notifications with Telegram alerts, and session management.
Password (argon2id) + TOTP (RFC-6238) SSO via a Caddy forward_auth checkpoint, no shared-domain cookie
Operator console: live health tiles, sign-in metrics, security notifications with Telegram alerts, per-session revoke
Live in production gating multiple Mochi apps, backed up nightly and monitored off-host
TypeScriptTOTPCaddy
Mochi Strata
Engineer · 2026
A fleet-wide database control plane
The fleet's database control plane: one binary, deployed per host with a declarative host profile, keeping every PostgreSQL cluster and SQLite database safe, reliable, fast, and lean. Covers drift, health, backup with drill-restore, bloat reclaim, and tuning — report-first, mutating only under an explicit --apply, with a baselined-ratchet metric so a capability that regresses or can't run reports RED instead of silently skipping.
One binary manages both PostgreSQL (WSL clusters + the adopted prod cluster) and the fleet's SQLite databases
Report-first by default: every capability plans and reports, mutating only under an explicit --apply
Drill capability reloads a backup into an ephemeral scratch cluster and verifies row-counts + checksums
TypeScriptPostgresSQLite
Lucid
Engineer · 2026
A local, offline text-to-speech studio
A local, fully offline text-to-speech studio built on VoxCPM2 — no cloud calls, no external API. A CLI, a PySide6 desktop GUI, and an MCP server all drive the same engine seam, backed by a persistent SQLite-backed project store and a render cache keyed by script, voice, and style. Verified end-to-end against real VoxCPM2 inference on an RTX 5090.
CLI, PySide6 GUI, and MCP server share one engine seam and one SQLite-backed project store
Verified end-to-end against real VoxCPM2 inference on an RTX 5090
A worker/IPC split isolates the model process from the app, with supervised restart and crash-safe reaping
PythonVoxCPM2PySide6MCP
AuspiceNode
Engineer · 2026
An autonomous quant trading agent for crypto
A solo-dev quant trading AI agent for Binance Spot — market-data ingest, rule-based/ML/autonomous-AI strategies, and a FastAPI + WebSocket backend behind a TradingView-chart web UI. AI strategies are first-class traders, but every signal, human or AI, clears the same risk gate, position caps, and kill switch before an order is placed. Pre-alpha, defaults to Binance testnet.
AI strategies are first-class traders — every signal clears the same risk gate, caps, and kill switch as rule-based/ML strategies
Decimal-native backtest engine (no float) with hyperopt, walk-forward, and Pareto analysis
FastAPI + WebSocket backend serving a TradingView Lightweight Charts web UI, defaults to Binance testnet