We Tried to Send 500M Notifications a Month on a $25 Server
Notification infrastructure has a funny cost curve. Sending one notification is cheap. Sending millions reliably is where it gets interesting, and most hosted platforms charge per message, so the bill grows with usage.
We wanted to know how much hardware the work itself needs. So we squeezed notifkit, our open-source notification service, plus Postgres and Redis into 2 vCPUs, the size of a $25/month cloud server, and load-tested it.
The first answer was 375 notifications a second, and we published it. Then we ran the test for 15 minutes instead of 10 seconds, with failing sends, retries and people opening emails, and the answer got smaller.
See the project on GitHubMeasured on a laptop in Docker with a simulated email provider, not on a cloud server with real providers. What this doesn't prove.
The setup
Everything ran in Docker with hard CPU and memory limits, so the whole stack fit in 2 vCPUs: one for notifkit, half each for Postgres and Redis. A load generator sent notifications to 100,000 users across 8 templates, and a fake provider stood in for email and SMS, taking 150 to 250 ms per send like a real one. Nothing was paid for or delivered to a real inbox.
Where it breaks
The first tests ran for 10 seconds with nothing going wrong, and said 375 a second. Real traffic is messier, so the second test kept everything running at once. 3% of sends failed and went through the real retry path (again after 30 seconds, then 2 minutes). 30% of delivered messages reported an open 20 seconds later, and 15% of those a click. Two simulated dashboard users read the delivery log every 2 seconds.
Then we held a fixed rate for 15 minutes and counted how many notifications were still waiting at the end. If the box keeps up, that number stays small. If it can't, it grows for the whole run.
Same $25 box, everything running. In and out are notifications per second. Waiting is everything queued but not yet sent. Median delay is API call to provider.
| In | Out | Waiting at 15 min | Median delay | Kept up |
|---|---|---|---|---|
| 200/s | 199.5/s | 145 | 0.4 s | yes |
| 400/s | 236.7/s | 144,247 | 3 min | no |
Why we tried 400/s: the one-minute ramp
Before the 15-minute runs, a quick ramp started at 150/s and went up 50/s every minute, to find a rate worth soaking. Each rate ran for only one minute, which is too short to see a slow pile-up: 400/s looked fine here and failed over 15 minutes.
| In | Out | Waiting at 1 min | Passed |
|---|---|---|---|
| 150/s | 148.5/s | 74 | yes |
| 200/s | 199.0/s | 170 | yes |
| 250/s | 249.5/s | 178 | yes |
| 300/s | 299.1/s | 0 | yes |
| 350/s | 345.4/s | 321 | yes |
| 400/s | 391.5/s | 478 | yes |
| 450/s | 365.1/s | 4,076 | no |
| 425/s | 345.4/s | 5,287 | no |
After 450/s failed, the ramp tried one step halfway back (425/s), which failed too, so 400/s went on to the 15-minute run.
Here's how the queue moved during both runs:
Notifications waiting, sampled about every 10 seconds. Flat means it keeps up.
- 400 a second
- 200 a second
Table view
| Into the soak | 400/s backlog | 200/s backlog |
|---|---|---|
| 1 min | 3,060 | 2,271 |
| 3 min | 15,284 | 89 |
| 6 min | 42,682 | 237 |
| 9 min | 75,593 | 202 |
| 12 min | 109,561 | 280 |
| 15 min | 144,247 | 145 |
The 400/s run went straight after the ramp, on a database that already held about 340 MB. The 200/s run started on an empty one.
At 400 a second the box delivered 237 a second on average, and only 205 a second in the second half. By the end 144,247 notifications were waiting, the median one took 3 minutes to reach the provider, and the API's p95 had climbed to 478 ms.
At 200 a second it kept up. A backlog of up to 3,354 built in the first minute and a half as the first retries and open webhooks arrived, then cleared, and stayed between 85 and 400 for the rest of the run. The median notification reached the provider in 0.4 s. p95 was 14 s and p99 41 s, because the 3% that failed waited for their retry. The API answered in 7 ms at the median and 121 ms at p95. Five notifications ended up in the dead-letter queue out of 179,987.
In this setup, the notification pipeline sustained about 200 notifications a second for 15 minutes without building a queue. That's 17 million a day.
We haven't found the exact limit between 200 and 400, and we haven't run anything longer than 15 minutes. Both are next.
What we got wrong the first time
- The test was too short375/s in 10 seconds, 400/s for one minute, and not even 400/s over 15 minutes. Retries, webhooks and a growing database don't show up in a 10-second test.
- Notifkit wasn't the busiest partDuring the failed 400/s soak, Postgres ran at 83% of its half-vCPU limit and Redis at 77%. The notifkit process sat at 64%, lower than the 85% it used at 200/s, which suggests it was waiting on the databases. We haven't profiled it far enough to say which one gave out first.
- A bigger box doesn't scale linearlyA single 3 vCPU node only kept 1.44 cores busy, barely more than a 2 vCPU node (1.35). More below.
Bigger box or more workers?
When one box isn't enough there are two options: make it bigger, or keep nodes small and add more of them. Notifkit can split into one API node
that takes requests plus 1 vCPU / 1 GB worker nodes that do the sending. Every node runs the same code and the services option
picks its role (deployment guide). We tried both, giving Postgres and Redis more room at each size.
These scaling runs only got the simple test: give the workers a backlog and time how fast they clear it, for 10 seconds, three runs per size. We didn't ramp or soak any of them, and nothing failed or retried. Read them as a comparison between setups, not as rates any of them will hold.
Every app node in green is 1 vCPU / 1 GB. Each group shares the same Postgres and Redis. 10-second catch-up test only, no soak. Median of three runs.
- 1 vCPU nodes: one API node plus workers
- One bigger node
Table view
| App nodes | Postgres | Redis | Catch-up | API call to provider |
|---|---|---|---|---|
| 1 node, 1 vCPU / 1 GB | 0.5 vCPU / 512 MB | 0.5 vCPU / 256 MB | 570/s | 2.0 s |
| 1 API + 1 worker, each 1 vCPU / 1 GB | 1 vCPU / 1 GB | 0.5 vCPU / 512 MB | 658/s | 0.6 s |
| 1 node, 2 vCPU / 2 GB | 1 vCPU / 1 GB | 0.5 vCPU / 512 MB | 1,078/s | 2.2 s |
| 1 API + 2 workers, each 1 vCPU / 1 GB | 1.5 vCPU / 1.5 GB | 1 vCPU / 1 GB | 1,031/s | 1.3 s |
| 1 node, 3 vCPU / 3 GB | 1.5 vCPU / 1.5 GB | 1 vCPU / 1 GB | 1,392/s | 2.5 s |
The bigger box wins on raw throughput. A single 2 vCPU node cleared 1,078/s; an API node plus one worker, same total CPU, cleared 658/s. At 3 vCPUs the single node reached 1,392/s against 1,031/s for an API plus two workers.
Workers win on latency and keep scaling. With the API separate from the worker, a notification reached the provider in 0.6 s, against 2.2 s on the single 2 vCPU node. And the single node stopped using the CPU it was given: 1.35 cores busy at 2 vCPUs, 1.44 at 3.
So we'd keep the API small and add 1 vCPU workers as volume grows, and give Postgres and Redis more room early, since the soak says they run out first. Given what the soak did to the $25 box's number, we'd expect every setup here to hold less than this over a long run.
What this benchmark does not prove
A benchmark is only useful if you know what it measured.
It doesn't prove that a $25 cloud server can send 520 million real emails in a month, and the original 970 million even less so. That figure multiplies a 15-minute rate by 30 days. Everything ran on a laptop in Docker, and the provider was simulated: each send took 150 to 250 ms, but nothing was delivered and no provider was paid.
A real deployment adds things this test didn't cover:
- provider rate limits (more on those next)
- real network latency, and provider outages longer than a single failed send
- a database that has been growing for weeks, not 15 minutes
- disk I/O on cloud volumes
- CPU-credit throttling on burstable instances
It's a measurement of the pipeline on this hardware. It isn't a production SLA.
Is 200 a second enough?
It depends on the channel. These are the production limits for transactional messages on an established account, not the starter limits a new account gets:
| Channel | Provider | Production limit |
|---|---|---|
| SMS | Twilio, A2P 10DLC | Up to 225 req/s across major US carriers, at the highest brand trust score |
| SMS | Twilio, verified toll-free | Up to 150+ req/s once approved for high throughput |
| SMS | Twilio, short code | From 100 req/s, more as a paid upgrade |
| Amazon SES | Set per account, raised as your volume and reputation grow; no published ceiling | |
| SendGrid | 10,000 req/s | |
| Push | Firebase Cloud Messaging | 600,000 a minute per project, about 10,000 req/s |
Checked October 2026. One request sends one message. Carriers actually count SMS segments: a text up to 160 characters is one, so a typical code or alert is one request. A longer text splits into several and uses up the limit faster.
SMS: the carriers cap you first. Even with the best trust score, one 10DLC brand gets about 225 req/s, and a short code starts at 100. A box that sustains 200 a second is already at or above what a typical SMS sender is allowed to send.
Email and push: the provider won't stop you, but the volume is huge. An established SES or SendGrid account can go well past 200 a second, and FCM allows 50 times that. 200 a second is 17 million transactional emails a day, though, every day: password resets, receipts, alerts. Few products send that much.
If you do need a provider's full rate limit, add workers. Each 1 vCPU worker runs the whole sending pipeline, so going faster means adding a node, not rebuilding anything (see bigger box or more workers). Give Postgres and Redis more room as you go, since they were the first to run hot in the soak.
Either way, peaks matter more than the average. When a provider rate-limits a send, notifkit parks it with the scheduler and tries again when the window reopens (architecture), up to its retry limit. At the 2,633/s the API accepted in our tests, a burst of 50,000 notifications is queued in about 20 seconds, then goes out as fast as your provider allows.
What it costs to run
The $25 pays for the server. Keep it busy and four more things show up on the bill: disk for the delivery log, backups of that disk, bandwidth out to your email and SMS providers, and CPU credits, because a $25 instance is burstable and only runs at 20% of each vCPU for free. Here is the whole bill for one AWS box keeping 30 days of logs, at AWS list prices (us-east-1, October 2026):
| Per month | 10M | 100M | 518M (200/s, all month) |
|---|---|---|---|
| Server: t4g.medium, 2 vCPU / 4 GB | $25 | $25 | $25 |
| CPU credits above the burst baseline | $0 | $0 | $31 |
| Disk: gp3, 30 days of logs | $3 (38 GB) | $16 (199 GB) | $76 (950 GB) |
| Backups: daily snapshots, kept a week | $1 | $11 | $58 |
| Bandwidth out: 10 KB a message, first 100 GB free | $0 | $81 | $458 |
| Total | $29 | $133 | $648 |
| Total sending through SES in the same region | $29 | $52 | $190 |
Bandwidth is the big one at high volume. We assumed 10 KB leaves the box per notification, about one HTML email sent to a provider's API. SMS and push messages are a tenth of that. If your email provider is Amazon SES in the same region as the box, that row drops to zero, as long as the box has a public IP; from a private subnet, a NAT gateway charges $0.045 per GB unless you add a VPC endpoint for SES. Provider fees per message are extra in every case.
These rows are calculated, not measured. Disk uses the 1.9 KB per notification the database grew by during the 200/s soak, including open and click events. CPU credits use the 1.47 cores the box burned at 200/s.
What running it yourself takes
- BackupsPostgres holds your users, templates and delivery log. Back it up like any production database.
- Log pruningThe database grew by about 1.9 KB per notification, around 18 GB a month at 10 million. Keep as much history as you need and delete the rest.
- Redis headroomRedis remembers each message for an hour so a retry is never sent twice. Size it for an hour of your peak traffic.
- UpgradesNew versions ship as an npm package; bump it and redeploy. Database migrations run when the server starts.
When self-hosting makes sense
At low volume, hosted services are simpler and often free: Knock, Courier and Novu all include 10,000 messages a month. Usage-based pricing gets substantial at high volume, and that's where running it yourself starts to be worth the work.
Hosted plans at their published per-message price, against one self-hosted box with everything on the bill from the table above. Provider fees for email, SMS and push are extra in every case.
- Knock
- Courier (dashed)
- Novu Cloud
- Self-hosted, one box
Table view
| Per month | Knock (Starter) | Courier (Business) | Novu Cloud (Team) | Self-hosted |
|---|---|---|---|---|
| 1M | $5,000 | $4,950 | $1,150 | $26 |
| 10M | $50,000 | $49,950 | $11,950 | $29 |
| 50M | $250,000 | $249,950 | $59,950 | $75 |
| 100M | $500,000 | $499,950 | $119,950 | $133 |
Prices are taken from each company's public pricing page, October 2026: Knock Starter, Courier Business, Novu Team, with the per-message overage extended to each volume. These are list prices. All three offer contracts at high volume, and a negotiated price can be much lower. Knock and Courier bill per message; Novu bills per workflow run. Knock and Courier differ by $50 at most, so their lines overlap.
Self-host notifkit if
- You send enough that per-message pricing hurts.
- Engineers own notifications. Templates and workflows are code in your repo, changed by pull request.
- You already run Postgres and Redis. Or you don't mind adding two well-known services.
- The data has to stay with you. Your users' details only leave your servers in the messages you send.
Use a hosted service if
- Non-engineers edit the messages. Novu, Knock and Courier come with a visual editor; notifkit doesn't.
- You want a drop-in in-app inbox. Notifkit can deliver to a webhook, but the UI is yours to build.
- You don't want to run servers. Or be the one paged when they break.
- Your volume is small. There's no bill to save.
Try it
If the left column sounds like you, the quickstart takes about ten minutes. You start a notifkit server next to Postgres and Redis, create a project, add a template and a user, and send your first notification. It's free and MIT-licensed, and your app can be written in any language: it only talks to notifkit over HTTP.
Reproduce it
The benchmark is in the notifkit repo under profiling/. Docker has to be running.
cd profiling
# 10-second tests: accept, catch up, keep up (3 runs per size)
A="--tiers=20,40,60 --ingest=10 --steady=10 --runs=3"
npm run profile:horizontal -- $A
npm run profile:vertical -- $A
# Ramp from 150/s in +50/s one-minute steps, then a 15-minute soak at the highest that passed
npm run profile:sustained
# 15-minute soak at a fixed rate
npm run profile:sustained -- --soak-rate=200
Exact configuration
notifkit 1 vCPU / 1 GB running every service (API, enricher, engine, delivery, scheduler, events), Postgres 0.5 vCPU / 512 MB, Redis 0.5 vCPU with AOF
(256 MB in the 10-second tests, 512 MB in the soaks). Worker concurrency 50, delivery concurrency 400, 5 Postgres connections.
NODE_ENV=production, LOG_LEVEL=info. Load comes from a k6 container on the same Docker network. 100,000 seeded users, 8 templates,
35% critical, 40% normal and 25% low priority.
Will you get these numbers on a real server?
Not exactly, and we can't say which way. A server could be faster: no Windows VM layer, and the load generator isn't competing for the same CPU. It could be slower: a cloud vCPU is often slower per thread than a laptop core boosting past 4 GHz, burstable instances throttle, and real providers rate-limit you. The six 10-second runs of the $25 box varied by 5% (583 ± 28/s catching up), mostly from background load on the laptop. Each soak ran once.