Case studyOur own product2026

Conductor: an AI project manager

Conductor is an agentic AI project manager for software teams. It plans the work, watches GitHub, reviews pull requests against your own coding standards and closes a task only when the code is actually merged. We designed and built it end to end, and now run it for teams as a single-tenant subscription.

The product lives at conductor.originsphere.in (opens in a new tab)

One isolated stack per customer, on our servers or yours, with the model provider you choose.

When does a task close?

Checked by code, not a model

  1. Condition 1: A linked pull request is merged into the default branch
  2. Condition 2: The latest code review covers that exact merged commit
  3. Condition 3: The review didn't request changes or flag anything high-severity
  4. Condition 4: The task's acceptance criteria are marked met

All four hold

The task closes, with the evidence attached.

Any one fails

It stays open, and the reasons get written down.

agents across three graphs
12
agents across three graphs
role-checked API endpoints
74
role-checked API endpoints
automated tests per pull request
960+
automated tests per pull request

An agentic AI project manager, in short

Conductor's agents read project documents and résumés and turn them into an assigned task plan. Then they watch GitHub, review pull requests against the team's own coding playbooks, close tasks when a fixed rule proves the work is merged, raise risk alerts and write status emails grounded in real records.

Anything with real consequences, like assignments, sensitive feedback or outbound email, waits for a person to approve it. Everything the agents do on their own lands in an audit log. And nothing is ever marked done because someone, or something, said it was.

We built all of it, from the agent architecture to the admin console. It now ships as a single-tenant deployment on a monthly subscription: one isolated stack per customer, on our infrastructure or yours, with your own model provider.

The problem: project management that trusts status updates

Most trackers end up recording what people said, not what shipped.

Software delivery runs on second-hand information. A project manager asks “is it done?”, a developer answers from memory, a ticket gets dragged to Done, and the board quietly becomes a record of what people said rather than what shipped. Reviews lag, coding standards drift from project to project, and somebody loses Friday afternoon stitching the weekly status report together from Slack, GitHub and email.

The first wave of “AI for project management” made this worse. A chat assistant that summarises a board still depends on the board being true. Ask a language model whether a task is complete and it will cheerfully say yes.

The brief we set ourselves

  1. 1

    Ground everything in real data

    Task status comes from commits, pull requests and review outcomes. Never from a language model's opinion.

  2. 2

    Keep people in charge of high-impact actions

    Assignments, sensitive feedback and outbound email pause until someone approves them.

  3. 3

    Make the agents the product

    The API and the console exist to configure, watch and overrule the agents. The world has enough CRUD trackers.

  4. 4

    Stay provider-agnostic

    Local inference today, a hosted model tomorrow. Switching is a configuration change, not a rewrite.

  5. 5

    Sell it as a private, single-tenant system

    The inputs are résumés, emails and source code: exactly the data a services company won't put in a shared SaaS.

How it works: narrow agents, boring routing, explicit gates

The intelligence is in how the agents are wired together, not in one giant chat loop. Each agent is a node in a LangGraph graph that makes a single structured LLM call and hands back data. Routing between nodes is plain Python, and “is it done?” is answered by a pure function over database rows.

Three graphs, one lifecycle

Planning, oversight and communication each get their own graph. Brass marks the places where a person decides.

Graph 1

Planning supervisor

Started by an admin. Human-gated.

  1. Trigger

    Upload the BRD or FRD and your developers' résumés

  2. AI step

    Profile agent

    Builds a skill profile for each developer.

  3. AI step

    Planner

    Breaks the documents into modules and tasks with acceptance criteria, estimates, priorities and dependencies.

  4. Human control

    You approve the task list

  5. AI step

    Assignment agent

    Proposes who does what from the profiles, the project's developer pools and a configurable active-task capacity.

  6. Human control

    You approve the assignments

  7. System update

    Provisioning agent

    Creates the GitHub repositories and webhooks.

Change orders go back through the same graph as a diff of added, modified and cancelled tasks.

Graph 2

Oversight pipeline

Started by GitHub. Autonomous, but overridable.

  1. Trigger

    A signed GitHub webhook event arrives

  2. AI step

    Commit monitoring

    Keeps track of pushes and pull requests as they land.

  3. AI step

    Review agent

    Reads the diff, the task's acceptance criteria and the active playbooks for the stacks it touches, then records a verdict with findings.

  4. Rules & checks

    The done rule

    Closes the task, or records exactly why it's blocked.

  5. Human control

    Admin override

    Force-complete or reopen, with a stated reason that lands in the audit log.

Graph 3

Communication graph

Started by an event or a schedule. Gated by kind of email.

  1. Trigger

    Review feedback, a risk alert or the weekly executive digest is due

  2. AI step

    Composer

    Writes the email from real records, not from memory.

  3. Human control

    Send or hold

    Goes out straight away or waits for approval, depending on the policy for that kind of email.

  4. System update

    Inbox

    Replies are read from the mailbox, threaded and embedded into memory.

  • Trigger
  • AI step
  • Rules & checks
  • Human control
  • System update

One rule decides “done”

The completion step doesn't ask a model anything. A task closes only when a linked pull request is merged into the repository's default branch, the latest review covers that exact merged commit, the review didn't request changes or flag a high-severity finding, and the acceptance criteria are marked met.

Miss any of that and the task stays open, with the blocking reasons recorded. An admin can still force-complete or reopen it, but they have to say why, and the reason goes into the audit log. A pull-request description that claims the criteria are met counts for nothing: third-party text is treated as data, never as evidence.

Memory that retrieves instead of retraining

“Train the AI on our standards and emails” is done with retrieval, not fine-tuning. Emails, reviews, completions and activated playbooks are embedded into PostgreSQL with pgvector and pulled back in before an agent writes or judges anything.

A Recon agent distils uploaded coding-standard documents into a structured rule set for each technology stack. Once an admin activates it, code review retrieves the matching playbook for every pull request and can request changes on a clear violation.

A daily risk monitor with no LLM in the loop

Every day, plain code works out where delivery is slipping:

  • Overdue tasks
  • Stale tasks
  • Blocked tasks
  • Developers over capacity
  • Tasks with no activity
  • Pull requests open too long
  • Milestones at risk of their target date

Only the alert email is written by a model. Repeat findings are latched for a configurable cooldown, so nobody gets paged twice for the same problem.

A console for watching and overruling the agents

The Next.js console opens on Today: decisions waiting for you, grounded risks, agents currently running, and a Measure Strip that shows each project's live status mix at a glance. A first-run checklist walks a new tenant from “model configured” to “first project planned” in eight checks.

Decisions
The human-in-the-loop queue for plans, assignments and gated emails.
Projects
Workspaces with overview, tasks, team, risks and updates tabs.
Workboard
Every task as a board, grouped list, table or timeline. You can only drag a card where an audited endpoint sits behind the move.
People
The roster, skill profiles built from résumés, and assignment pools.
Playbooks
Per-stack coding standards, from upload through distillation to activation.
Mail, Activity and AI usage
Email threads with reply-gap triage, the full audit log, and token usage per agent.
Settings
Change the model provider, GitHub, SMTP or IMAP at runtime, with inline credential checks and secrets encrypted at rest.

Why single-tenant, and why a subscription

Résumés, private repositories and internal email don't belong in a shared database.

Conductor handles résumés, private repositories, code diffs and internal email. A shared multi-tenant database is the wrong shape for that data, and for the people who own it. So single-tenant packaging was a design constraint from day one, and it's now the commercial model too.

Every subscription gets its own stack

  • Its own PostgreSQL database and encryption keys
  • Its own API and console containers
  • Its own GitHub app credentials and mailbox
  • Its own model endpoint

Nothing is pooled between customers.

Runs anywhere Docker runs

The stack is three containers behind a reverse proxy. Local inference through LM Studio keeps every prompt inside your network, and switching to OpenAI is a single setting. Microsoft Entra ID single sign-on and Microsoft Graph mail are there for Microsoft 365 organisations.

What the subscription covers

  • Provisioning your tenant on our infrastructure, or a guided install on your own server.
  • Versioned releases from a private registry, each pinned to the exact build that passed the test pipeline, with one-command rollback.
  • Setting up the GitHub, email and model integrations, plus your first playbooks.
  • Security patches, migrations and monitoring, with support hours and an uptime commitment defined per plan.
  • Plans priced per tenant and sized by active projects and developers, so there's no per-seat maths.

Want numbers? Ask us for current plans and we'll size one to your projects and team.

Engineering facts

What went into the build, counted from the repository rather than estimated.

What was delivered in Conductor, by area
AreaWhat was delivered
Agents12 agent modules across 3 LangGraph graphs, a deterministic supervisor, and a change-order path that reuses the planning graph
API74 endpoints across 21 routers, with role-based access control on every route (admin, executive, developer)
Data16 Alembic migrations; PostgreSQL 16 with pgvector for memory; secrets encrypted at rest with Fernet
IntegrationsGitHub webhooks with HMAC verification plus polling, SMTP and IMAP, Microsoft Graph mail, Microsoft Entra ID SSO, and LM Studio or OpenAI through one factory
Audit54 distinct audit event types covering every autonomous action and every manual override, each with actor, timestamp and reasoning
Tests685 backend tests and 277 frontend tests, run on every pull request before an image is published
CodebaseAbout 18,000 lines of typed Python and 20,500 lines of TypeScript
DeliveryPull-request-gated pipeline: test, publish to a container registry, deploy pinned by commit, serialised rollouts

Security posture

The console hides what you can't use. The backend is what actually says no.

  • Role-based access is enforced by the backend on every endpoint. The console's guards are there for convenience only.
  • JWT access tokens live in memory, with rotating refresh tokens, reuse detection and revocation on password change.
  • Login, refresh and SSO exchange are rate-limited, and failed logins are audited.
  • GitHub webhook signatures are verified in constant time, payload size is capped before parsing, and every document upload route has a size cap.
  • In production the app refuses to boot with default or weak secrets, and hides the API docs.
  • Prompts treat third-party text as data. A pull-request description that claims the acceptance criteria are met is never taken as evidence.

What we learned building it

Four things we'd tell anyone building agents that touch real systems.

  1. Deterministic beats agentic wherever you can write the rule.

    The completion decision, the risk monitor and the graph routing are all plain functions. They're testable, a customer can follow them, and prompt drift can't touch them.

  2. Side effects around an approval pause need discipline.

    A graph node that pauses for approval runs again from the top when it resumes, so every database write lives in a node that runs exactly once. That's where the three-node “propose, await, apply” pattern behind every gate came from.

  3. Idempotency is the integration layer's whole job.

    Webhooks and polls both write with upserts, reviews are keyed by pull request and commit, and inbound mail is deduplicated on message id. A replay never creates a duplicate task, review or email.

  4. Telemetry is a callback, not a code change.

    Token usage per agent, flow and project comes from a single LangChain callback attached where every run is configured.

Conductor: questions people ask

What is an agentic AI project manager?

Software in which AI agents do the project-management work themselves: planning tasks, assigning them, reviewing code and reporting status. That's different from an assistant that answers questions about a board people keep up to date. Conductor's agents do the work and pause for a human decision at the points that matter.

How does Conductor decide a task is complete?

With a fixed rule, not a model. A linked pull request must be merged into the default branch, the latest code review must cover that merged commit, the review must not have requested changes or found a high-severity issue, and the acceptance criteria must be marked met. Otherwise the task stays open and the blocking reasons are recorded.

Which AI models does Conductor use?

Any OpenAI-compatible endpoint. Tenants typically run a local model through LM Studio so prompts and code never leave their network, or use OpenAI by changing one setting. The provider is configured, never hard-coded.

Is Conductor multi-tenant?

No. Each subscription is an isolated single-tenant deployment with its own database, encryption keys, credentials and model endpoint.

Can we host it ourselves?

Yes. The stack is three Docker containers behind a reverse proxy. OriginSphere provides the images, the install and the updates as part of the subscription.

Does it replace our project tracker?

It replaces the parts that depend on people typing status into a tracker. Tasks, assignments and completion live in Conductor and are grounded in GitHub. Developers keep working in Git and pull requests as usual.

What integrations are supported today?

GitHub (webhooks and polling), SMTP and IMAP email, Microsoft Graph mail, Microsoft Entra ID single sign-on, and LM Studio or OpenAI for inference.

How is developer data protected?

Profiles and emails are treated as sensitive. They're used to assist and inform, never to surveil. Access is role-based and enforced by the backend, developers see only their own data, secrets are encrypted at rest, and every autonomous action can be audited and overridden.

Want Conductor running on your own GitHub?

Book a walkthrough of a live tenant, or ask for a pilot on your own repositories. We reply within one business day.

We reply within one business day, and we're happy to sign an NDA first. Prefer email? Write to info@originsphere.in.

Chat with us