Engineering notes
Boring where it counts.
You are about to hand a tool read access to your Microsoft estate. That deserves more than a logo wall, so here is what GraphPaaS is actually made of, how a sync runs end to end, and what happens when Microsoft's APIs push back. No architecture-astronaut diagrams — just the choices and the reasons behind them.
The stack
Every choice, and why
Nothing exotic. Each layer was picked because the problem asked for it, and each one is boring enough to still be running in five years.
Async by default, so a request waiting on Microsoft Graph never blocks a worker. Pydantic gives typed request/response contracts and OpenAPI docs for free.
Tenant syncs are long, bursty and rate-limited. They belong in a queue, not in an HTTP request. Priority lanes keep a manual sync ahead of the scheduled ones.
Relational data only: organizations, tenants, users, plans, sync jobs, audit log. Schema changes ship as reviewed Alembic migrations.
Collected metrics are columnar snapshots, not rows to update. Parquet reads a 90-day trend in one shot, and erasure for a tenant is a prefix delete.
Server components render the dashboard shell; only the interactive parts ship JavaScript. TypeScript strict end to end.
Handles time-series with gaps, drill-down interactions and a light/dark palette swap without a chart rewrite per theme.
You sign in with the Microsoft account you already have. Tenant access is granted by standard admin consent — no local passwords for customer accounts.
Every service is a container built from the same repo, deployed as one versioned unit. Everything runs inside the EU.
Host and container metrics are scraped by Prometheus; logs ship to Loki and both land in Grafana, where the dashboards and alerts live. A disk filling up or a worker that died is something we get paged about — not something a customer has to report.
How a sync runs
From consent to chart
Schedule
A scheduler fans out one job per connected tenant, per domain. Each lands on a priority queue: an admin-triggered sync jumps ahead of the scheduled batch.
Collect
A worker acquires an application token for that tenant, then runs read-only collectors against Microsoft Graph and Azure billing APIs — each one a pure function of the token, so it is testable without a live tenant.
Store
Results are written as a dated Parquet snapshot under that tenant's own prefix. Yesterday's snapshot is never overwritten — that is what makes trends possible.
Serve
The API reads the requested snapshot, applies the plan's entitlements and returns only that tenant's rows. The dashboard renders; no query ever touches another tenant's data.
When things go wrong
Built for the bad day
Graph APIs throttle, permissions get revoked, licences expire. The interesting engineering is not the happy path.
Throttle before they throttle you
Microsoft publishes per-tenant request limits, and crossing them gets you slowed down for everyone. Our HTTP client counts requests per tenant in a shared window and backs off proactively — rather than discovering the ceiling by hitting it.
Retry that respects the server
A 429 or a transient 5xx is retried with exponential backoff, honouring the Retry-After header the API sends. A 401 or 403 is not retried — revoked consent is a decision, not a blip, and hammering it helps nobody.
Degrade, don't disconnect
Some reports need a licence tier or a permission a tenant hasn't granted. Those collectors return a marker explaining what's missing instead of failing the whole sync — one gated report never costs you the other nine.
Every run is a record
Each sync writes a job row: domain, status, duration, rows collected, error. That history is what turns "the number looks odd" into an answer, and it is what the alerting below is built on.
We tell you before you notice
A dashboard quietly serving three-day-old data is worse than one that is obviously down. If a tenant stops syncing you get a banner on the dashboard and an email to your admins — the day after consent is withdrawn, or after three days of any other failure. Stale data always says so.
Tenant isolation
Six gates, one request
“Each tenant is isolated” is what every vendor says. Here is the actual sequence a single call goes through before one row of your data is returned — and where it stops if any of it fails.
- 01
The call arrives
Every metrics request carries a credential: either the session token issued when you signed in with Microsoft, or an API key you generated yourself. No credential, no route.
✕ Missing or malformed → 401
- 02
Who is asking
The token's signature is verified, then the account behind it is loaded fresh from the database on every single call. A deleted or expired account is rejected even while its token is still cryptographically valid — the token is a claim, not the answer.
✕ Deleted, expired or revoked → 401
- 03
Which tenant, and may you
Name no tenant and you get your own. Name one and it is looked up under your organisation specifically — the check is not "does this tenant exist", it is "is this tenant yours". An MSP reaches all of its clients this way; nobody reaches past their own organisation.
✕ Tenant outside your organisation → 403
- 04
Are you entitled to it
The organisation's plan is checked for that specific domain, and point-in-time reads are clamped to the plan's history window. Entitlement is resolved server-side from the plan — the client has no say in it.
✕ Domain not on your plan → locked
- 05
Where the data lives
The snapshot's location is not built from your request. It is read off a stored job record that was already filtered to the tenant approved in step 03 — so the address of the data is a consequence of the authorisation, not an input to it.
- 06
What comes back
That one snapshot is read and returned. There is no larger result set being narrowed down, which means there is no filter to accidentally omit.
Step 05 is the one that matters. Most cross-tenant leaks in this class of product are a filter somebody forgot to add. Here there is nothing to forget: a request never names a location, it names a tenant — and by then the tenant has already been checked.
Security & data
Your estate, not ours
GraphPaaS holds a read-only mirror of data your organisation already owns. Everything below exists to keep it that way.
The Microsoft permissions we request are read scopes. GraphPaaS reports on your estate; it does not change it. Anything that would write is a separate, explicitly-consented decision.
Access tokens and customer-supplied API keys are encrypted before they are stored, and cached with a TTL shorter than their own lifetime. Nothing sensitive is written in plaintext.
You can ask for a specific tenant — you cannot ask for one that isn't your organisation's, and the location your data is read from is never assembled from anything you sent. Walked through step by step below.
User principal names, display names and IP addresses are hashed or omitted in application logs. Debugging a sync should not mean reading a customer's staff directory.
Processing and storage stay inside the EU. We act as data processor under a DPA; retention is configurable per tenant and erasure removes both the database rows and the stored snapshots.
Removing a user anonymises their identifying fields rather than orphaning references — so an erasure request is honoured without corrupting the audit trail it belongs to.
The full terms are in our privacy policy and terms. Doing security review for a client and need something specific — a DPA, a permissions breakdown, a sub-processor list? Ask us directly.
The part that looks easy
“How hard can it be?”
Every feature on this page started life as a one-line ticket. None of them stayed that way. A sample, so you know the depth is earned:
“Just call the Graph API.”
It pages, it throttles, and the same endpoint answers differently depending on which licence tier the tenant holds. Half the work is deciding what to show when the answer is "your licence doesn't include that".
“Just list who never signs in.”
A tenant can switch on concealed names, and the usage reports come back with anonymised identifiers instead of people. Your per-user table is decorative until an admin turns that setting off — so we detect it and say so, rather than showing a table of nobody.
“Just remove the unused licences.”
An assignment can be direct, inherited from a group, or both at once. Remove the wrong one and it reappears at the next group sync, or it takes someone's mailbox with it. This is why we surface waste and hand you the link, instead of clicking it for you.
“Just show the Azure spend.”
Two different billing APIs, two different shapes, and access that is not granted by the same consent as everything else — a missing billing role comes back as a 404 rather than a 403, which is a genuinely enjoyable afternoon to debug.
“Just refresh it every few hours.”
Microsoft's usage reports trail real time by a couple of days. Labelling that data "today" would make every number on the page a small lie, so the dashboard tells you which day it is actually showing you.
None of it is hard, exactly. There is just an enormous amount of it — and every line above was learned the same way: in production, on a real tenant, at an hour nobody would choose. You get the version that already knows.
Who builds it
Built by an admin, for admins
GraphPaaS started as a set of scripts for the reports Microsoft makes you assemble by hand. It is a small, focused product — which means the person who writes the collector is the person who answers your email about it. Every feature ships with tests, and the whole suite runs before anything reaches production.