Triamorph Systems

← Engineering Dispatches / Full-Stack

B2B SaaS Billing Architecture: How to Implement Multi-Tier Subscriptions, Usage Metering, and Stripe Webhook Idempotency

By Hammad Haider · 13 min read read

Building a robust billing engine for B2B SaaS is notoriously prone to edge cases: double billing from duplicate webhook deliveries, race conditions during mid-cycle plan upgrades, prorated seat expansions, and dropped usage metrics. As software enterprises scale from seat-based pricing to hybrid consumption-based billing models, relying solely on basic Stripe client-side SDKs is insufficient. This architectural dispatch demonstrates how to design an enterprise-grade billing state machine, persist an immutable audit ledger in PostgreSQL, and enforce guaranteed exactly-once webhook processing.

Architectural Takeaways

  • Store customer billing state as an immutable state machine in PostgreSQL, treating Stripe as a payment processor rather than the single source of truth for access entitlement.
  • Enforce atomic idempotency using a dedicated `stripe_webhook_events` PostgreSQL table with advisory row locks to prevent duplicate credit provisioning during network retries.
  • Decouple real-time API request ingestion from Stripe Metering API calls using TimescaleDB hyper-tables and asynchronous batch workers.

1. Relational Billing Data Model & Subscription State Machine

Many SaaS startups make the critical architectural mistake of querying Stripe directly on every protected API call to check if an organization has an active plan. This introduces external network latency (300ms+ per request) and causes complete application outages whenever Stripe encounters degraded performance.

An enterprise billing architecture stores customer subscription states, feature quotas, and seat limits in local PostgreSQL tables with Row-Level Security (RLS). Stripe customer IDs and subscription IDs are stored as foreign keys. When a user authenticates, their tenant permissions are resolved locally in under 1ms.

2. Zero-Loss Stripe Webhook Processing with Idempotency Keys

Stripe webhooks use an "at-least-once" delivery guarantee. If your server takes more than a few seconds to respond with a 200 OK, or if a temporary network blip occurs, Stripe will retry sending the exact same webhook payload up to dozens of times over 72 hours.

Without strict idempotency guards, handling an `invoice.payment_succeeded` or `checkout.session.completed` event multiple times will provision duplicate seats, send repetitive email receipts, or grant double balance credits. We implement atomic lock tables with cryptographic signature verification.

3. Real-Time Consumption Metering & Aggregation Engine

In usage-based software (such as AI token consumption, compute hours, or active data storage), piping every single event synchronously to Stripe’s Metering API creates rate-limiting bottlenecks and high egress costs.

The optimal architecture ingests usage events into high-throughput Redis stream buffers, aggregates them periodically into TimescaleDB or clickstream tables, and dispatches summarized hourly or daily usage records to Stripe using the Stripe Metered Billing API.

4. Handling Mid-Cycle Upgrades, Prorations & Grace Periods

When an enterprise client adds 20 team members midway through their annual subscription, Stripe automatically calculates line-item prorations. Your backend must seamlessly update the `allocated_seats` constraint and trigger immediate seat provisioning without disrupting active user sessions.

Similarly, when payment attempts fail, transitioning subscriptions to `past_due` and initiating a 7-day dunning sequence with automated banner warnings preserves revenue while avoiding abrupt tenant churn.

Read more technical guides on our Dispatches Index →