jobsearch v0.0.1

← scaleai / AI Product Manager (Coding/Multimodal)

cover_letter / art_KAaJ1ENEJFs

role
scaleai / AI Product Manager (Coding/Multimodal)
model
anthropic/claude-sonnet-4.6
created
2026-09-22T16:26

↓ Download .docx

Cover letter

Dear Scale AI Hiring Team, Scale AI sits at the infrastructure layer of the AI stack — the place where frontier model quality is actually determined, not just claimed. That specificity matters to me. My path from hand-coding backpropagation through time in C++ at UC Berkeley in 2004, to publishing at NeurIPS, to building a production RL post-training workbench that benchmarks GRPO, DPO, PPO, and nine other algorithms across TRL, VeRL, OpenRLHF, and NeMo RL today, has been defined by a conviction that the quality of training data and evaluation infrastructure is the real lever on model capability. Scale's mission to build reliable AI systems for the world's most important decisions is precisely the problem I want to work on. **Technical and AI Foundation** My AI/ML work is not adjacent to my product work — it is the same work. In 2025–2026 I built an RL post-training workbench from scratch: a three-phase platform covering Reward Lab (designing and A/B testing reward functions across RLVR, learned, and hybrid configurations on GSM8K, MATH, HumanEval, and UltraFeedback), a live training Playground running real TRL-powered GRPO and DPO jobs with SSE metric streaming on Apple Silicon MPS and CUDA, and an Arena for head-to-head framework benchmarking with GPU passthrough in Docker containers. Implementing 12 RL algorithms — PPO, GRPO, DAPO, REINFORCE, REINFORCE++, RLOO, DPO, SimPO, IPO, KTO, ORPO, SPPO — and standardizing throughput, memory, and convergence benchmarking across four frameworks gave me a working understanding of the tradeoffs that matter when you are trying to improve model behavior through post-training data. Parallel to that, I built aeval, a local-first model evaluation platform with five core eval types (factuality, reasoning, instruction-following, safety, code generation), adversarial safety testing with refusal detection, and statistical rigor via bootstrap confidence intervals, Welch's t-test, Cohen's d effect size, and saturation detection — with CI/CD regression gates. The stack runs FastAPI, TimescaleDB, Redis, Next.js, and Ollama. These are not toy projects; they are the infrastructure I use to make decisions about model behavior, and they map directly to the data specification, quality review, and evaluation workflow responsibilities in this role. On the multimodal side, I repurposed StreamIO's screen capture and HLS streaming pipeline to build AutoEval, an automated visual evaluation system for robot model training that scores grasp poses, segmentation maps, and bounding boxes against natural-language rubrics using Claude and GPT-4V — reducing evaluation cycles from 72 hours to approximately four minutes. Zero-integration screen capture means it works against any visualization tool without instrumentation. That architecture — multimodal AI as an evaluation layer over arbitrary visual outputs — is directly relevant to Scale's image, video, and world model verticals. My NeurIPS 2014 paper on neural networks for protein secondary structure prediction, and the 2026 PyTorch rewrite of that system spanning 413 parameters to 8 billion (a 19-million-fold scale increase), anchor the research credibility behind the product work. **Why This Role** My career arc has moved from deep technical implementation through platform-scale product management and back into hands-on AI building. What I have not yet done is apply that combination inside an organization whose core product *is* AI data quality at frontier scale. The Scale AI PM role for Coding and Multimodal data verticals is the specific intersection I have been building toward: RL data products, agentic systems, coding benchmarks, and multimodal evaluation are exactly the domains where my workbench and eval platform work is most directly applicable. The JD's emphasis on driving execution across Coding, Agentic, and RL data products, coordinating cross-functional delivery, and contributing to data specification and evaluation workflows describes work I have done — but at Scale the stakes and the feedback loop are orders of magnitude larger. The opportunity to shape the data that trains frontier models, rather than consuming those models as a product builder, is the specific pull. **Selected Prior Experience** - Built a 3-phase RL post-training workbench implementing 12 algorithms (PPO, GRPO, DAPO, DPO, and eight others) with cross-framework benchmarking (TRL, VeRL, OpenRLHF, NeMo RL), live SSE metric streaming, and GPU Docker passthrough — directly applicable to Scale's RL data product work. - Built aeval, a model evaluation platform with adversarial safety testing, refusal detection, data contamination detection via SHA-256 hashing, and statistical significance tooling (bootstrap CI, Welch's t-test, Cohen's d) — applicable to Scale's evaluation workflow and data quality responsibilities. - Architected OpenClaw multi-agent orchestration framework (gateway protocol, subagent delegation, profile management, session switching) coordinating AI agent workflows across multiple industries — relevant to Scale's Agentic data vertical. - At Intuit, delivered ICE Self-Service platform reducing developer onboarding from 2–3 weeks to minutes, achieved 275% YoY growth in ICE engagements scaling to 675M+ in FY23, and scaled throughput from 6K to 50K TPS — demonstrating cross-functional program execution and platform delivery at scale. - Owned Search Service, Search Catalog, and SPL/SPL2 at Splunk; delivered Scheduler Service end-to-end in approximately four months and led a query performance initiative achieving up to 10x improvements for a beta Fortune 500 customer — demonstrating milestone-driven delivery against customer commitments. - Built Vantage's AI interview prep platform with 40+ agent tools, async job polling, idempotent endpoints, and SSE streaming over FastAPI — demonstrating full-stack AI product execution including coding-adjacent tooling. - NeurIPS 2014 published researcher; UC Berkeley B.S. in Computational Engineering Science with Lawrence Berkeley National Laboratory computational biology research — establishing the research foundation behind the product work. **Closing** Scale's mission — reliable AI systems for the world's most important decisions — requires that the data powering those systems be held to the same standard. My background building RL training infrastructure, model evaluation platforms, and multimodal AI pipelines, combined with 12+ years of cross-functional platform product delivery, positions me to contribute immediately to the Coding and Multimodal data verticals while growing into full ownership of product and customer outcomes. I would welcome the opportunity to discuss how my work maps to Scale's current priorities. Sincerely, **O. Felix Amoruwa** famoruwa@berkeley.edu · 909-731-9011 · felixamoruwa.info