Autotelica AI · Remote first role. Preference for UK-based hybrid, with occasional attendance at the client’s London office · target part-time start: December 2026, target full-time start: February 2027
Autotelica is building a platform that lets a large UK FTSE client's developers use AI coding agents under the company's own standards. The core is a Claude Code plugin: an MCP server, slash commands and hooks that load the right rules into the agent's context, let the agent test its own work against them, and write the evidence into each commit. A pull-request gate checks the same rules again before anything merges.
You own that core, and you're the engineer the client sees every day. You work to the platform's technical authority and lead the squad's engineering.
What you'll do
- Own the context and hooks design: session start, pre- and post-tool-use, commit time, pull request.
- Port the hooks from Claude Code to GitHub Copilot, then to other agent harnesses.
- Build and run the pull-request gate and the evidence record that ships with every change.
- Keep the rule checks honest: deterministic where a regex will do, LLM-judged against a rubric where it won't, and handed to a named human where neither can decide.
- Own security and standards compliance of the platform itself: threat model, pen-test fixes, secrets handling.
- Sit with the client's engineering teams, answer their questions, and turn what you hear into backlog items.
- Review the squad's code and set its engineering standards.
What you bring
- Eight or more years building production software, with at least two leading a team.
- Deep, daily use of Claude Code, Copilot, Cursor or Codex. You've written hooks, skills, MCP servers or plugins, not just used them.
- A clear view of what an agent must never do unchecked, and how you'd stop it.
- TypeScript or Python to a high standard. Git internals, CI and GitHub Actions.
- Ability to be be in client London office (paddington) when needed.
Nice to have
- Secure SDLC, static analysis, or policy-as-code tools such as OPA or Semgrep.
- LLM evaluation: rubrics, graders, test fixtures for model output.
#J-18808-Ljbffr