Applied AI Engineer - Part time
We are hiring an Applied AI Engineer to build EXCEED’s AI functionality end to end. The center of the role is building agents — agentic features that plan, use tools, and act across the platform — surrounded by the full range of applied AI work: RAG pipelines, the content-generation pipeline, hands-on labs, and new product features from front end to model call.
Your daily coding instrument is Claude Code — we are an AI-native engineering team — but everything you build must be model-agnostic by design: the functionality you ship supports Claude, OpenAI, and Gemini alike. You will think in provider-portable abstractions, open standards (MCP), and evaluation methods that hold across models, so that a new provider or model generation is a configuration change, not a rebuild.
What You Will Do
- Build agents — the core of the role. Agentic features across EXCEED: agents that plan, call tools, hold memory, and act within guardrails — orchestrated with LangChain/LangGraph, exposed through MCP where it fits, and running against Claude, OpenAI (GPT), and Gemini behind a provider-abstraction layer with model routing and cost management.
- Ship end-to-end product features. From React front end through Node services and MongoDB to the model call on Google Cloud — you own features whole, not layers of them.
- Build and tune RAG. Ingestion, chunking, retrieval quality, reranking, and grounded generation with citations — the retrieval backbone under content generation and in-product assistance.
- Evolve the content-generation pipeline. Grounded generation from ingested sources, structured outputs that survive provider differences, and accuracy validation so generated modules, quizzes, and labs stay true to their sources.
- Build the labs. Hands-on lab environments in our sandboxed execution platform: mock APIs, auto-verification harnesses that grade learner builds, and managed multi-provider model access with per-learner budgets.
- Own evaluation infrastructure. Task datasets, deterministic and model-based graders, regression suites, and cross-provider benchmarks — for our features and for the models we route between.