Joy Dong
AI enablement & agent operations · New York, NY
I run multi-vendor AI agent fleets in production, every day: Anthropic, OpenAI, xAI and local models working under one memory layer, with evaluation gold standards, adversarial review gates, spend caps, and kill switches. Then I teach other people to do it.
Agent-assisted building is my declared working mode. I architect, review, evaluate, and operate the systems. I do not present myself as a hand-coding engineer, and I do not need to be one to make an agent fleet safe, measurable, and cheap to run.
What I Run
A production agent fleet, in daily use, not a demo environment.
Multi-vendor fleet under one memory layer
Agents across two machines and three model vendors share a single corrections ledger, promoted pattern cards, and per-session learning journals, so any agent picks up any task with the previous agent's lessons already loaded.
Guardrails that are enforced, not aspirational
At-most-once publishing state machines for anything that leaves the machine, programmable per-action-class rate caps and spend budgets, kill-switch hierarchies, and evidence gates that block publication when a claim cannot be traced to a live source.
Evaluation and labeling operations
A TREC-style gold standard for a social platform's search stack: 140 queries and 1,017 graded pairs, frozen under a versioned protocol with 14 numbered grading rules, dual-reviewer rounds, provenance isolation between rating-side and engine-side artifacts, and a hard freeze gate.
Cross-model adversarial review
Nothing substantive ships without a review gauntlet that runs parallel Anthropic reviewers against an OpenAI adversarial pass. Logged cases of one model family catching a defect the other missed are past forty.
Vendor-agnostic by construction
A 70-plus skill library kept in sync across vendors by an idempotent engine. Every operational value (caps, budgets, windows, model choice) changes through config, never through a code edit.
Autonomous account agents, taken live
A two-account autonomous posting system on a real social platform: the full watch, evidence, draft, gate, publish, digest loop with evidence checking, per-account budgets, operator digests, and idempotent recovery. When the platform was retired the architecture was extracted into eight reusable vendor-neutral skills.
Newsletter automation with a watchdog
Daily TEA ships every weekday through an automated pipeline: draft, cover render, schedule, verify, dedup ledger, plus a watchdog that catches silent non-publication. One incident in its history, root-caused and gated the same week.
Proof of Work
Three artifacts being sanitized out of the private infrastructure above and published under my name. Dates are commitments, not aspirations. No link appears here until the thing behind it is real.
→
at-most-once publishing state machine
repo publishing September 2026
The exactly-once discipline behind every agent action that leaves the machine: durable intent records, idempotency keys, crash-safe recovery, and the reasoning about which failures are safe to retry.
→
search / recommendation eval gold-standard toolkit
repo publishing September 2026
The grading protocol, rater rules, adjudication path, and freeze gate from the evaluation lab above, generalized into a toolkit anyone can run against their own retrieval or recommendation stack.
→
MCP server
repo publishing October 2026
A working Model Context Protocol server for a real workflow, built agent-assisted and small enough that I can whiteboard its internals cold.
Case Studies
Full customer-shaped builds, written up with the decisions, the tradeoffs, and what I would do differently.
→
Outreach agency in a box
write-up publishing October 2026
Inbound request through triage, tracker, personalized drafts, and follow-up scheduling, with an eval harness and a decision memo. Built against my own institutional workload, then generalized.
→
Compliance-shaped document extraction with verification
write-up publishing November 2026
Extraction under a regulated-industry bar, where an unverified answer is worse than no answer. Verification layer, refusal behavior, and the cost of getting it wrong.
Build Recordings
Ninety-minute timed builds from a cold prompt, screen-recorded and narrated, so you can watch how I actually work with agents rather than take my word for it. The best three get published.
→
Selected recordings
publishing November 2026
Recording weekly through the autumn. Nothing is published here until it is worth your time to watch.
How I Work
Agent-assisted building is the declared mode, not a disclaimer. I direct agents to write the code; I own the architecture, the review, the evaluation, and the operational discipline that decides whether it ships. That is the job I am applying for, and it is the job I already do every day.
What that buys you: someone who has already hit the failure modes that only appear in production. Silent non-publication. Duplicate posts from a retried job. A green test suite in front of a dead user-facing surface. A model confidently citing a source it never read. Every one of those has a written correction and a gate behind it in my systems.
What it does not buy you: a hand-coding engineer. If the role's core is writing production code unassisted at a whiteboard, I am the wrong candidate and I will say so on the first call.
Background
Employer Relations Specialist
Columbia University, SEAS · New York, NY · 2025–present
Run employer recruitment programs for an Ivy League engineering school, and built the agent-assisted operations layer for the role itself: email triage protocol, job-posting pipeline with a human gate (drafts only, never auto-send), outreach tracking, scheduled agent check-ins. A live case study in AI enablement inside a large institution.
Executive Operations Manager, digital tech venture
Private holding company · New York, NY · 2021–2024
Co-led the launch of a regulated stablecoin ecosystem in six months. Evaluated 30-plus vendors, managed 20 contracts worth $10M, cut monthly burn by half, and stood up operations hubs adopted by ten teams.
Educator
Avenues · Columbia · New Oriental · 2014–2022
Eight years teaching complex material to non-expert audiences, from K-5 bilingual through adult ESL. Enablement is a teaching job before it is a technical one.
Credentials & Stack
Education · M.A. Applied Linguistics/TESOL, Columbia University Teachers College · B.Eng. Computer Science, Nankai University · Advanced AI Product Management & Leadership, Maven
Agent ops · multi-agent orchestration · MCP · agent memory design · eval design (gold standards, LLM-as-judge limits, inter-rater protocols) · guardrails (caps, kill switches, at-most-once delivery) · cost and latency management · prompt-injection defense · cross-model review
Stack fluency · Anthropic and OpenAI SDKs · tool use · prompt caching · RAG and embeddings (Ollama with sqlite-vec, in production) · Modal · Supabase/Postgres · GitHub Actions · launchd and cron operations
Languages · English and Mandarin, professional in both