🧠 AI Tooling

AI tools that
actually ship.

I build the tooling that makes AI output usable at ship quality: an LLM gateway with routing and spend metering, an asset studio that holds one art style across a catalog, and agents that run real work. ~15 years of Unity UI and shipped live products is what tells me when the output is wrong.

Luke Litman headshot
69
Repos Built
6
AI Systems Built
32
Live Deployments
54
Active in 90 Days

"I evaluate AI output because I know what correct looks like. When a model generates code, designs, or analysis, I can tell whether it's right, not because I ran a test suite, but because I've spent ~15 years building the same kinds of Unity UI and game systems by hand."

(On AI Evaluation)
Case Study

How I evaluate AI-generated work

A compact example of the rubric I use when reviewing AI output: not just whether it looks plausible, but whether it would survive ship.

🧪

Scenario: Game UI Generated by AI

An AI model produces a mobile game upgrade screen: hero art, currency balances, CTA hierarchy, reward copy, and layout annotations. My review separates surface polish from ship readiness.

UX Review Visual QA Game Systems Implementation Notes

Evaluation Criteria

  • Task fit: Does the screen solve the actual player/job-to-be-done?
  • Information hierarchy: Are currency, cost, reward, and next action immediately legible?
  • Ship feasibility: Can the design be implemented with reusable components and sane asset budgets?
  • Brand consistency: Does the output match tone, visual system, and platform constraints?
  • Failure modes: What would break under localization, small screens, missing data, or live-ops variants?
Why the judgment holds up

Domain Depth

  • ~15 years of Unity UI and software architecture, I know when generated code will break at scale
  • Professional art director, I know when generated visuals miss the brief
  • Shipped game developer, I know when game logic is subtly wrong
  • Product leader, I know when a feature spec has gaps

Multi-Model Judgment

  • I use 5+ models daily and know each one's strengths and failure modes
  • I can select the right model for the right task, not just default to one
  • I architect prompts as systems, not one-off queries
  • I evaluate outputs against real-world standards, not just "does it look right"

"The final score is never a vibe check. I return specific pass/fail notes, severity, suggested fixes, and the reason each issue matters to players, developers, or the business."

(AI Evaluation Workflow)
Capabilities

What I actually build with AI

🤖

Agent Systems

Assistants that do more than answer. They take actions inside the product, carry memory between sessions, and read the page or account they are sitting on. Built for a fintech app, a game platform, and a grants product.

🔗

LLM Infrastructure

One gateway in front of several providers, with auth, routing, and spend metered per call. Cost stays visible per service instead of scattered across API keys, and services never hold provider keys directly.

🛑

Human in the Loop

Approval gates on anything that writes or spends. A run will hold for as long as the decision takes, and what happened gets written back somewhere a person can read it later. Nothing auto-approves.

🎯

Evaluation

The part that still needs a person: knowing when generated code, art, or analysis is wrong. Having shipped the same work by hand for ~15 years is what makes that call quick instead of theoretical.

🖼️

Generative Art Pipelines

Generative image workflows folded into real creative pipelines rather than run as one-offs, and held to art direction standards before anything ships. Art direction is the day job, so the bar does not move.

Daily Toolkit

Models & platforms

OpenAI Ecosystem

GPT-4o GPT-4 o1 o3 DALL-E 3 Assistants API Function Calling Embeddings Vision

Other Models

Claude (Anthropic) Grok (xAI) Gemini (Google) Midjourney

Infrastructure & Tools

Vercel GitHub Python JavaScript / TypeScript REST APIs Slack Integrations Linear Google Workspace
Character Stats

Spec Sheet

Hands-on daily with live AI systems. Every rating comes from real deployment experience.

🧠 AI & Prompt Engineering

Prompt Engineering
Master
Agent Architecture
Advanced
Multi-Model Orchestration
Advanced
AI Output Evaluation
Master
Generative AI (Images)
Advanced
RAG / Embeddings
Proficient

🤖 AI Platforms & Models

OpenAI (GPT-4o, o1, o3)
Master
DALL-E 3
Advanced
Claude (Anthropic)
Advanced
Midjourney
Proficient
Gemini (Google)
Familiar
Grok (xAI)
Familiar
Vibe Familiar Proficient Advanced Master
Selected Projects

AI systems in ship

💳

INK Pay : In-App AI Assistants

Conversational UI · In-App Actions · Support Agent Fleet

Several of the assistants living inside INK's fintech and social app. They answered questions about the company and the product, walked users through the app, and took real actions in it: navigating, changing color schemes, and the rest of what a financial and social app assistant is expected to handle. Behind them, a separate fleet of agents in the operations console handling customer support and working through user issues.

💬

Maiden (GM) : Assistant Rail

Agent Systems · Persistent Memory · Page Context

A shared assistant rail for the studio's sites. Threads, persistent memory, and a page-context hook that lets the assistant read the page it is sitting on. Personas load from a shared skill library, and every model call routes through Gate Forge instead of a direct provider key.

🔑

Gate Forge : Identity and LLM Gateway

LLM Gateway · Auth · Spend Metering

One session and one chat completions endpoint sitting in front of multiple providers. Every call is authenticated and metered, so spend is visible per service rather than scattered across API keys.

🔎

Muse : Open Call Assistant

Conversational Matching · Curated Data · Human Ops Desk

The assistant inside Open Call, which helps artists find the support that already exists for them. You talk about the work, the city, and the month you are in; Muse matches against a curated database of grants, health coverage, and workforce programs. A human ops desk keeps the data honest.

📝

Brief Light : Agent-Driven Interview

Agent Interview · Human Approval · Gamejam

Inside Gamejam, an agent interviews you about a game idea and turns the conversation into a structured brief and design wiki. Every write and every spend waits on an explicit human yes.

Node Forge : Human-Gated Workflows

Workflow Graphs · Human in the Loop · Scheduled Runs

Graphs that chain studio tools together and stop at the steps that need a person. Fired by hand, webhook, or schedule, they will hold a run paused for as long as the decision takes, then write the record of what happened back to a repository.

🤖

WAP : Multi-Persona Agent Platform

AI Architecture · OpenAI API · Agent Design

Custom-built AI platform with multiple specialized personas, each with distinct system prompts, capabilities, and evaluation criteria. Orchestrates complex workflows across different AI models.

⚒️

IronReach : AI-Powered Brand Studio

Agent Orchestration · Live AI · Multi-Client

Runs two autonomous AI agents across workspaces, coordinating brand development, project management, and infrastructure for 5+ simultaneous client engagements at ironreach.xyz. Live client work, running daily.

Need an AI architect who
actually understands the output?

I bring two decades of cross-domain expertise to AI evaluation, system design, and agent architecture.

Download Resume → Contact Me