The cost of AI coding

Copilot is not your whole AI coding bill.

Short answer

The real cost of AI coding is spread across at least four invoices: GitHub Copilot seats and credits, direct LLM API spend on providers like Anthropic and OpenAI, model usage on Vertex AI, Azure OpenAI or Bedrock, and the CI minutes agents consume when they run. Judging AI coding cost from the Copilot line alone understates it, often substantially.

Four places AI coding spend hides

They rarely arrive on the same invoice, and often not to the same budget owner:

GitHub Copilot
Seats plus metered credits. The visible one, and usually the one that gets managed because it is the one with a name attached.
Direct LLM API keys
Anthropic, OpenAI and others, billed to a card or an account that engineering opened during a spike. Frequently invisible to finance until renewal.
Cloud-hosted models
Vertex AI, Azure OpenAI, Bedrock. Buried inside a cloud bill that is already large enough that a five-figure line does not stand out.
CI and infrastructure
Agent runs consume Actions minutes and pull data. The compute cost of an AI workflow lands on a completely different meter from the tokens.

Why the total matters more than any one tool

Optimising Copilot while an unmonitored API key doubles quarter over quarter is a rounding-error exercise. The question worth answering is not 'what does Copilot cost' but 'what does AI-assisted engineering cost us, and is it moving the way we intended'.

That total is also the only honest basis for a build-versus-buy conversation. Teams routinely conclude that calling an API directly is cheaper than a seat licence, having compared a metered rate against a subscription without counting the engineering time, the CI minutes, or the credits already included in the seat they are still paying for.

Cost per active developer is the comparable unit

Totals across four invoices grow whenever the company does, which makes them impossible to judge. Divide by the developers actually using any of it and you get a figure that stays meaningful across reorganisations, hiring waves and tool changes.

It also makes the awkward comparison possible: what a developer costs you in AI tooling, against what that tooling is expected to return. That comparison is uncomfortable precisely because it is the right one.

The ROI question, framed honestly

AI coding ROI claims tend to fail on the denominator, not the numerator. Time saved is genuinely hard to measure and easy to overstate; cost is knowable and usually understated by three of the four sources above.

The defensible version is narrower and more useful: track cost precisely, track a small number of delivery signals consistently, and report the two side by side without asserting a causal link the data cannot support. A leader given an accurate cost and an honest set of delivery trends will make a better decision than one given a confident productivity multiplier.

Bringing it into one view

FinOpsAid reads GitHub's meters and, where you connect them, your Anthropic, Vertex AI, Azure OpenAI and Bedrock usage — reconciling each against its own provider's billing rather than blending them into a single invented total. Figures that come from a billing API are labelled billed; figures computed by applying a rate to observed usage are labelled modeled, with the date they were computed.

FinOpsAid connects through a GitHub App with read-only scopes. It changes no seats, edits no permissions, pushes no code, and never stores your source. Because only Read is granted, writing back to GitHub is not a policy we promise — it is technically impossible.

FAQ

Frequently asked questions

What does AI coding actually cost per developer?

It depends entirely on surface mix, but the mistake is scoping the question to Copilot. A complete figure adds Copilot seats and credits, direct LLM API spend, cloud-hosted model usage on Vertex, Azure OpenAI or Bedrock, and the CI minutes agent runs consume. Three of those four are usually invisible to whoever owns the Copilot budget.

Is calling an LLM API directly cheaper than Copilot?

Rarely, once you count everything. Direct API comparisons typically weigh a metered token rate against a seat subscription while ignoring the engineering time to build and maintain the integration, the CI minutes it consumes, and the credit allowance already included in the seat you are still paying for.

How do we track AI spend across GitHub and cloud providers?

Pull each provider's own usage and billing data and keep them separate, reconciling each against its own invoice rather than merging them into one blended figure. A combined total is useful for budgeting; a blended per-unit rate is not, because the meters count different things at different rates.

Can you measure AI coding ROI?

You can measure the cost precisely and the benefit only approximately, so report them side by side rather than as a ratio. Track spend accurately, track a small consistent set of delivery signals, and resist asserting causation the data cannot support. An honest cost with honest trends beats a confident productivity multiplier.

Try it

See these numbers for your own org

Connect your organization read-only and FinOpsAid meters it nightly — seats, credits, minutes, Codespaces and storage, attributed to teams and reconciled against your billing API. Every modeled figure is labelled an estimate, with the date it was computed.