jobsearch v0.0.1

← togetherai / Product Manager, AI Infrastructure

cover_letter / art_ESrHzfW_Stk

role
togetherai / Product Manager, AI Infrastructure
model
anthropic/claude-sonnet-4.6
created
2026-09-22T15:11

↓ Download .docx

Cover letter

Dear Together AI Hiring Team, Together AI is building the infrastructure layer that the AI ecosystem actually runs on — inference at 400+ trillion tokens a month, fine-tuning, RL post-training, and now the GPU clusters, storage, and observability that make all of it possible. That mission sits at the exact intersection where I've spent the last several years: not just shipping AI products, but building the platforms and tooling that let engineers move faster. When I scaled Intuit's ICE platform to 675M+ engagements in FY23 and drove throughput from 6K to 50K TPS via rSocket migration, I wasn't managing a roadmap handed to me — I was finding the architectural bottlenecks, instrumenting the data, and driving resolution across engineering, design, and infrastructure teams simultaneously. That's the operating mode this role describes, and it's the one I default to. **Technical and AI/ML Foundation** My technical foundation spans both the infrastructure layer and the AI stack above it. At Intuit, I owned developer-facing platform infrastructure across ~20 mobile apps and 30+ product SKUs — working directly in SQL and BigQuery to surface developer pain points, leading a GCP-to-AWS migration for Mailchimp MSaaS (Golang template, MySQL persistence, updated DevPortal), and building a Java JAR library to scan Git repos for configuration drift as part of the MSaaS Drift Detection program. I understand what it means to instrument a distributed system, read the telemetry yourself, and translate that into prioritized engineering work. On the AI side, I built a production RL post-training workbench from scratch in 2026 — three phases covering the full RLHF/DPO pipeline, with real TRL-powered GRPO/DPO training running on Apple Silicon (MPS) and CUDA, live SSE metric streaming, and a benchmarking arena for head-to-head framework comparisons across TRL, VeRL, OpenRLHF, and NeMo RL with GPU passthrough in Docker containers. I implemented 12 RL algorithms (PPO, GRPO, DAPO, REINFORCE, REINFORCE++, RLOO, DPO, SimPO, IPO, KTO, ORPO, SPPO) with standardized throughput, memory, and convergence benchmarking across frameworks. This is not adjacent experience to Together's infrastructure — it's direct experience with the workloads your GPU clusters are running. I also built aeval, a local-first model evaluation platform with a FastAPI orchestrator, TimescaleDB, Redis job queue, and Ollama integration — with bootstrap confidence intervals, Welch's t-test, Cohen's d effect size, and automated safety gates wired into CI/CD. Statistical rigor in data collection and hypothesis-driven experimentation isn't a methodology I've adopted for interviews; it's how I've built evaluation infrastructure. **Why This Role** Most PM roles at this stage offer either technical depth or product breadth — Together's AI Infrastructure PM role offers both, with a clear path to full ownership of observability or storage within nine months. The early focus on observability and GPU clusters maps directly to where I've built: I've instrumented distributed systems at scale, worked with GPU training workloads hands-on, and operated in the high-ambiguity, fast-moving environment that AI-native customers like Cursor, ElevenLabs, and Decagon live in every day. **Role-Specific Connection** The observability surface particularly resonates — at Intuit, I built and shipped the Asterias declarative asset lifecycle management platform with a GraphQL API, and used telemetry and usage data continuously to prioritize work across a large microservice surface. I know what it takes to instrument a system well enough to answer open questions, not just monitor known metrics. The GPU cluster workload is equally familiar: my RL workbench required reasoning about GPU memory, throughput, and framework-level performance tradeoffs across CUDA and MPS backends — the same reasoning your customers apply when choosing cluster configurations for training runs. **Selected Prior Experience** - Scaled Intuit's ICE platform to 675M+ engagements in FY23 (275% YoY growth); drove throughput from 6K to 50K TPS via rSocket migration supporting ~1.5M concurrent connections at sub-25ms TP99. - Led Mailchimp GCP-to-AWS migration for MSaaS: delivered Golang template, MySQL persistence integration, and updated DevPortal documentation to meet production deadline. - Built Java JAR library for MSaaS Drift Detection — scanned Git repos for configuration drift, partnered with Design on DevPortal UI, and built remediation roadmap using OpenRewrite. - Worked directly in SQL and BigQuery to surface developer pain points across ~20 mobile apps and 30+ product SKUs; built Asterias, a declarative asset lifecycle management platform with GraphQL API. - Built RL post-training workbench with GPU Docker passthrough, benchmarking TRL, VeRL, OpenRLHF, and NeMo RL across 12 algorithms with standardized throughput/memory/convergence metrics on CUDA and MPS. - Built aeval evaluation platform: FastAPI orchestrator, TimescaleDB, Redis job queue, Ollama — with bootstrap CIs, Welch's t-test, Cohen's d, and CI/CD regression detection. - Conducted enterprise-wide Service Language Assessment across 9 languages (Java, Python, Kotlin, Go, TypeScript, Scala, PHP, C++, Groovy), analyzing usage data and developer feedback to inform strategic decisions presented to CTO. - Delivered ICE Self-Service platform (DevPortal, GitOps config, ICE Playground), reducing developer onboarding from 2–3 weeks to minutes in pre-prod and under 24 hours for production, mitigating $1M+ in projected opex growth. **Closing** Together AI's mission — purpose-built infrastructure for AI engineers — is exactly the surface I want to own. The AI ecosystem is moving faster than general-purpose cloud can keep up with, and the teams building on Together need infrastructure that's instrumented, observable, and tuned for their workloads. I've spent 12 years finding the problems others haven't spotted yet in developer platforms and infrastructure, and the last two building hands-on AI systems from RL training pipelines to evaluation frameworks. I'd welcome the opportunity to bring that combination to the Together Cloud team. Sincerely, **O. Felix Amoruwa** famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info