Engineering notes

Boring where it counts.

You are about to hand a tool read access to your Microsoft estate. That deserves more than a logo wall, so here is what GraphPaaS is actually made of, how a sync runs end to end, and what happens when Microsoft's APIs push back. No architecture-astronaut diagrams — just the choices and the reasons behind them.

● Read-only Graph scopes● EU processing & storage● Per-tenant isolation

The stack

Every choice, and why

Nothing exotic. Each layer was picked because the problem asked for it, and each one is boring enough to still be running in five years.

API
FastAPI · Python 3.12

Async by default, so a request waiting on Microsoft Graph never blocks a worker. Pydantic gives typed request/response contracts and OpenAPI docs for free.

Background work
Celery · Redis

Tenant syncs are long, bursty and rate-limited. They belong in a queue, not in an HTTP request. Priority lanes keep a manual sync ahead of the scheduled ones.

Database
PostgreSQL 16

Relational data only: organizations, tenants, users, plans, sync jobs, audit log. Schema changes ship as reviewed Alembic migrations.

Metric storage
Parquet on S3-compatible object storage

Collected metrics are columnar snapshots, not rows to update. Parquet reads a 90-day trend in one shot, and erasure for a tenant is a prefix delete.

Frontend
Next.js 14 · App Router

Server components render the dashboard shell; only the interactive parts ship JavaScript. TypeScript strict end to end.

Charts
ECharts

Handles time-series with gaps, drill-down interactions and a light/dark palette swap without a chart rewrite per theme.

Identity
Microsoft OAuth2 · MSAL

You sign in with the Microsoft account you already have. Tenant access is granted by standard admin consent — no local passwords for customer accounts.

Runtime
Docker · EU infrastructure

Every service is a container built from the same repo, deployed as one versioned unit. Everything runs inside the EU.

Monitoring
Prometheus · Loki · Grafana

Host and container metrics are scraped by Prometheus; logs ship to Loki and both land in Grafana, where the dashboards and alerts live. A disk filling up or a worker that died is something we get paged about — not something a customer has to report.

How a sync runs

From consent to chart

01

Schedule

A scheduler fans out one job per connected tenant, per domain. Each lands on a priority queue: an admin-triggered sync jumps ahead of the scheduled batch.

02

Collect

A worker acquires an application token for that tenant, then runs read-only collectors against Microsoft Graph and Azure billing APIs — each one a pure function of the token, so it is testable without a live tenant.

03

Store

Results are written as a dated Parquet snapshot under that tenant's own prefix. Yesterday's snapshot is never overwritten — that is what makes trends possible.

04

Serve

The API reads the requested snapshot, applies the plan's entitlements and returns only that tenant's rows. The dashboard renders; no query ever touches another tenant's data.

When things go wrong

Built for the bad day

Graph APIs throttle, permissions get revoked, licences expire. The interesting engineering is not the happy path.

Throttle before they throttle you

Microsoft publishes per-tenant request limits, and crossing them gets you slowed down for everyone. Our HTTP client counts requests per tenant in a shared window and backs off proactively — rather than discovering the ceiling by hitting it.

Retry that respects the server

A 429 or a transient 5xx is retried with exponential backoff, honouring the Retry-After header the API sends. A 401 or 403 is not retried — revoked consent is a decision, not a blip, and hammering it helps nobody.

Degrade, don't disconnect

Some reports need a licence tier or a permission a tenant hasn't granted. Those collectors return a marker explaining what's missing instead of failing the whole sync — one gated report never costs you the other nine.

Every run is a record

Each sync writes a job row: domain, status, duration, rows collected, error. That history is what turns "the number looks odd" into an answer, and it is what the alerting below is built on.

We tell you before you notice

A dashboard quietly serving three-day-old data is worse than one that is obviously down. If a tenant stops syncing you get a banner on the dashboard and an email to your admins — the day after consent is withdrawn, or after three days of any other failure. Stale data always says so.

Tenant isolation

Six gates, one request

“Each tenant is isolated” is what every vendor says. Here is the actual sequence a single call goes through before one row of your data is returned — and where it stops if any of it fails.

  1. 01

    The call arrives

    Every metrics request carries a credential: either the session token issued when you signed in with Microsoft, or an API key you generated yourself. No credential, no route.

    ✕ Missing or malformed → 401

  2. 02

    Who is asking

    The token's signature is verified, then the account behind it is loaded fresh from the database on every single call. A deleted or expired account is rejected even while its token is still cryptographically valid — the token is a claim, not the answer.

    ✕ Deleted, expired or revoked → 401

  3. 03

    Which tenant, and may you

    Name no tenant and you get your own. Name one and it is looked up under your organisation specifically — the check is not "does this tenant exist", it is "is this tenant yours". An MSP reaches all of its clients this way; nobody reaches past their own organisation.

    ✕ Tenant outside your organisation → 403

  4. 04

    Are you entitled to it

    The organisation's plan is checked for that specific domain, and point-in-time reads are clamped to the plan's history window. Entitlement is resolved server-side from the plan — the client has no say in it.

    ✕ Domain not on your plan → locked

  5. 05

    Where the data lives

    The snapshot's location is not built from your request. It is read off a stored job record that was already filtered to the tenant approved in step 03 — so the address of the data is a consequence of the authorisation, not an input to it.

  6. 06

    What comes back

    That one snapshot is read and returned. There is no larger result set being narrowed down, which means there is no filter to accidentally omit.

Step 05 is the one that matters. Most cross-tenant leaks in this class of product are a filter somebody forgot to add. Here there is nothing to forget: a request never names a location, it names a tenant — and by then the tenant has already been checked.

Security & data

Your estate, not ours

GraphPaaS holds a read-only mirror of data your organisation already owns. Everything below exists to keep it that way.

Read-only by design

The Microsoft permissions we request are read scopes. GraphPaaS reports on your estate; it does not change it. Anything that would write is a separate, explicitly-consented decision.

Tokens encrypted at rest

Access tokens and customer-supplied API keys are encrypted before they are stored, and cached with a TTL shorter than their own lifetime. Nothing sensitive is written in plaintext.

Isolation is structural

You can ask for a specific tenant — you cannot ask for one that isn't your organisation's, and the location your data is read from is never assembled from anything you sent. Walked through step by step below.

PII stays out of logs

User principal names, display names and IP addresses are hashed or omitted in application logs. Debugging a sync should not mean reading a customer's staff directory.

EU, end to end

Processing and storage stay inside the EU. We act as data processor under a DPA; retention is configurable per tenant and erasure removes both the database rows and the stored snapshots.

Soft delete, real anonymisation

Removing a user anonymises their identifying fields rather than orphaning references — so an erasure request is honoured without corrupting the audit trail it belongs to.

The full terms are in our privacy policy and terms. Doing security review for a client and need something specific — a DPA, a permissions breakdown, a sub-processor list? Ask us directly.

The part that looks easy

“How hard can it be?”

Every feature on this page started life as a one-line ticket. None of them stayed that way. A sample, so you know the depth is earned:

“Just call the Graph API.”

It pages, it throttles, and the same endpoint answers differently depending on which licence tier the tenant holds. Half the work is deciding what to show when the answer is "your licence doesn't include that".

“Just list who never signs in.”

A tenant can switch on concealed names, and the usage reports come back with anonymised identifiers instead of people. Your per-user table is decorative until an admin turns that setting off — so we detect it and say so, rather than showing a table of nobody.

“Just remove the unused licences.”

An assignment can be direct, inherited from a group, or both at once. Remove the wrong one and it reappears at the next group sync, or it takes someone's mailbox with it. This is why we surface waste and hand you the link, instead of clicking it for you.

“Just show the Azure spend.”

Two different billing APIs, two different shapes, and access that is not granted by the same consent as everything else — a missing billing role comes back as a 404 rather than a 403, which is a genuinely enjoyable afternoon to debug.

“Just refresh it every few hours.”

Microsoft's usage reports trail real time by a couple of days. Labelling that data "today" would make every number on the page a small lie, so the dashboard tells you which day it is actually showing you.

None of it is hard, exactly. There is just an enormous amount of it — and every line above was learned the same way: in production, on a real tenant, at an hour nobody would choose. You get the version that already knows.

Who builds it

Built by an admin, for admins

GraphPaaS started as a set of scripts for the reports Microsoft makes you assemble by hand. It is a small, focused product — which means the person who writes the collector is the person who answers your email about it. Every feature ships with tests, and the whole suite runs before anything reaches production.