← All articles
Benchmark · October 2026

We Tried to Send 500M Notifications a Month on a $25 Server

Notification infrastructure has a funny cost curve. Sending one notification is cheap. Sending millions reliably is where it gets interesting, and most hosted platforms charge per message, so the bill grows with usage.

We wanted to know how much hardware the work itself needs. So we squeezed notifkit, our open-source notification service, plus Postgres and Redis into 2 vCPUs, the size of a $25/month cloud server, and load-tested it.

The first answer was 375 notifications a second, and we published it. Then we ran the test for 15 minutes instead of 10 seconds, with failing sends, retries and people opening emails, and the answer got smaller.

See the project on GitHub
200/ssustained: delivered nonstop for 15 minutes with failures, retries, opens and dashboard readsthe rate it can hold for a long time
2,633/saccepted by the API in a burst, queued and sent out afterwardsmedian of six 10-second runs
≈ 520Mnotifications a month at the sustained rate200/s × 30 days, if it held all month

Measured on a laptop in Docker with a simulated email provider, not on a cloud server with real providers. What this doesn't prove.

The setup

Everything ran in Docker with hard CPU and memory limits, so the whole stack fit in 2 vCPUs: one for notifkit, half each for Postgres and Redis. A load generator sent notifications to 100,000 users across 8 templates, and a fake provider stood in for email and SMS, taking 150 to 250 ms per send like a real one. Nothing was paid for or delivered to a real inbox.

the $25 box: 2 vCPU in totalYour appload generatornotifkit1 vCPU · 1 GBRedis0.5 vCPU · 256 MBPostgres0.5 vCPU · 512 MBEmail / SMSfake, 150–250 msAPI call

Where it breaks

The first tests ran for 10 seconds with nothing going wrong, and said 375 a second. Real traffic is messier, so the second test kept everything running at once. 3% of sends failed and went through the real retry path (again after 30 seconds, then 2 minutes). 30% of delivered messages reported an open 20 seconds later, and 15% of those a click. Two simulated dashboard users read the delivery log every 2 seconds.

Then we held a fixed rate for 15 minutes and counted how many notifications were still waiting at the end. If the box keeps up, that number stays small. If it can't, it grows for the whole run.

Fig 1Two 15-minute runs

Same $25 box, everything running. In and out are notifications per second. Waiting is everything queued but not yet sent. Median delay is API call to provider.

InOutWaiting at 15 minMedian delayKept up
200/s199.5/s1450.4 syes
400/s236.7/s144,2473 minno
Why we tried 400/s: the one-minute ramp

Before the 15-minute runs, a quick ramp started at 150/s and went up 50/s every minute, to find a rate worth soaking. Each rate ran for only one minute, which is too short to see a slow pile-up: 400/s looked fine here and failed over 15 minutes.

InOutWaiting at 1 minPassed
150/s148.5/s74yes
200/s199.0/s170yes
250/s249.5/s178yes
300/s299.1/s0yes
350/s345.4/s321yes
400/s391.5/s478yes
450/s365.1/s4,076no
425/s345.4/s5,287no

After 450/s failed, the ramp tried one step halfway back (425/s), which failed too, so 400/s went on to the 15-minute run.

Here's how the queue moved during both runs:

Fig 2Notifications waiting, over the two 15-minute runs

Notifications waiting, sampled about every 10 seconds. Flat means it keeps up.

  • 400 a second
  • 200 a second
050K100K150K0 min3 min6 min9 min12 min15 min400/s: 144,247200/s: 145
Table view
Into the soak400/s backlog200/s backlog
1 min3,0602,271
3 min15,28489
6 min42,682237
9 min75,593202
12 min109,561280
15 min144,247145

The 400/s run went straight after the ramp, on a database that already held about 340 MB. The 200/s run started on an empty one.

At 400 a second the box delivered 237 a second on average, and only 205 a second in the second half. By the end 144,247 notifications were waiting, the median one took 3 minutes to reach the provider, and the API's p95 had climbed to 478 ms.

At 200 a second it kept up. A backlog of up to 3,354 built in the first minute and a half as the first retries and open webhooks arrived, then cleared, and stayed between 85 and 400 for the rest of the run. The median notification reached the provider in 0.4 s. p95 was 14 s and p99 41 s, because the 3% that failed waited for their retry. The API answered in 7 ms at the median and 121 ms at p95. Five notifications ended up in the dead-letter queue out of 179,987.

In this setup, the notification pipeline sustained about 200 notifications a second for 15 minutes without building a queue. That's 17 million a day.

We haven't found the exact limit between 200 and 400, and we haven't run anything longer than 15 minutes. Both are next.

What we got wrong the first time

  • The test was too short375/s in 10 seconds, 400/s for one minute, and not even 400/s over 15 minutes. Retries, webhooks and a growing database don't show up in a 10-second test.
  • Notifkit wasn't the busiest partDuring the failed 400/s soak, Postgres ran at 83% of its half-vCPU limit and Redis at 77%. The notifkit process sat at 64%, lower than the 85% it used at 200/s, which suggests it was waiting on the databases. We haven't profiled it far enough to say which one gave out first.
  • A bigger box doesn't scale linearlyA single 3 vCPU node only kept 1.44 cores busy, barely more than a 2 vCPU node (1.35). More below.

Bigger box or more workers?

When one box isn't enough there are two options: make it bigger, or keep nodes small and add more of them. Notifkit can split into one API node that takes requests plus 1 vCPU / 1 GB worker nodes that do the sending. Every node runs the same code and the services option picks its role (deployment guide). We tried both, giving Postgres and Redis more room at each size.

These scaling runs only got the simple test: give the workers a backlog and time how fast they clear it, for 10 seconds, three runs per size. We didn't ramp or soak any of them, and nothing failed or retried. Read them as a comparison between setups, not as rates any of them will hold.

Fig 3Adding 1 vCPU workers, or one bigger node

Every app node in green is 1 vCPU / 1 GB. Each group shares the same Postgres and Redis. 10-second catch-up test only, no soak. Median of three runs.

  • 1 vCPU nodes: one API node plus workers
  • One bigger node
notifications sent per second, clearing a backlog (higher is better)Postgres 0.5 vCPU / 512 MB · Redis 0.5 vCPU / 256 MB1 node, 1 vCPU / 1 GB (the $25 box)570/sPostgres 1 vCPU / 1 GB · Redis 0.5 vCPU / 512 MB1 API + 1 worker, each 1 vCPU / 1 GB658/s1 node, 2 vCPU / 2 GB1,078/sPostgres 1.5 vCPU / 1.5 GB · Redis 1 vCPU / 1 GB1 API + 2 workers, each 1 vCPU / 1 GB1,031/s1 node, 3 vCPU / 3 GB1,392/s04008001,2001,600
Table view
App nodesPostgresRedisCatch-upAPI call to provider
1 node, 1 vCPU / 1 GB0.5 vCPU / 512 MB0.5 vCPU / 256 MB570/s2.0 s
1 API + 1 worker, each 1 vCPU / 1 GB1 vCPU / 1 GB0.5 vCPU / 512 MB658/s0.6 s
1 node, 2 vCPU / 2 GB1 vCPU / 1 GB0.5 vCPU / 512 MB1,078/s2.2 s
1 API + 2 workers, each 1 vCPU / 1 GB1.5 vCPU / 1.5 GB1 vCPU / 1 GB1,031/s1.3 s
1 node, 3 vCPU / 3 GB1.5 vCPU / 1.5 GB1 vCPU / 1 GB1,392/s2.5 s

The bigger box wins on raw throughput. A single 2 vCPU node cleared 1,078/s; an API node plus one worker, same total CPU, cleared 658/s. At 3 vCPUs the single node reached 1,392/s against 1,031/s for an API plus two workers.

Workers win on latency and keep scaling. With the API separate from the worker, a notification reached the provider in 0.6 s, against 2.2 s on the single 2 vCPU node. And the single node stopped using the CPU it was given: 1.35 cores busy at 2 vCPUs, 1.44 at 3.

So we'd keep the API small and add 1 vCPU workers as volume grows, and give Postgres and Redis more room early, since the soak says they run out first. Given what the soak did to the $25 box's number, we'd expect every setup here to hold less than this over a long run.

What this benchmark does not prove

A benchmark is only useful if you know what it measured.

It doesn't prove that a $25 cloud server can send 520 million real emails in a month, and the original 970 million even less so. That figure multiplies a 15-minute rate by 30 days. Everything ran on a laptop in Docker, and the provider was simulated: each send took 150 to 250 ms, but nothing was delivered and no provider was paid.

A real deployment adds things this test didn't cover:

  • provider rate limits (more on those next)
  • real network latency, and provider outages longer than a single failed send
  • a database that has been growing for weeks, not 15 minutes
  • disk I/O on cloud volumes
  • CPU-credit throttling on burstable instances

It's a measurement of the pipeline on this hardware. It isn't a production SLA.

Is 200 a second enough?

It depends on the channel. These are the production limits for transactional messages on an established account, not the starter limits a new account gets:

ChannelProviderProduction limit
SMSTwilio, A2P 10DLCUp to 225 req/s across major US carriers, at the highest brand trust score
SMSTwilio, verified toll-freeUp to 150+ req/s once approved for high throughput
SMSTwilio, short codeFrom 100 req/s, more as a paid upgrade
EmailAmazon SESSet per account, raised as your volume and reputation grow; no published ceiling
EmailSendGrid10,000 req/s
PushFirebase Cloud Messaging600,000 a minute per project, about 10,000 req/s

Checked October 2026. One request sends one message. Carriers actually count SMS segments: a text up to 160 characters is one, so a typical code or alert is one request. A longer text splits into several and uses up the limit faster.

SMS: the carriers cap you first. Even with the best trust score, one 10DLC brand gets about 225 req/s, and a short code starts at 100. A box that sustains 200 a second is already at or above what a typical SMS sender is allowed to send.

Email and push: the provider won't stop you, but the volume is huge. An established SES or SendGrid account can go well past 200 a second, and FCM allows 50 times that. 200 a second is 17 million transactional emails a day, though, every day: password resets, receipts, alerts. Few products send that much.

If you do need a provider's full rate limit, add workers. Each 1 vCPU worker runs the whole sending pipeline, so going faster means adding a node, not rebuilding anything (see bigger box or more workers). Give Postgres and Redis more room as you go, since they were the first to run hot in the soak.

Either way, peaks matter more than the average. When a provider rate-limits a send, notifkit parks it with the scheduler and tries again when the window reopens (architecture), up to its retry limit. At the 2,633/s the API accepted in our tests, a burst of 50,000 notifications is queued in about 20 seconds, then goes out as fast as your provider allows.

What it costs to run

The $25 pays for the server. Keep it busy and four more things show up on the bill: disk for the delivery log, backups of that disk, bandwidth out to your email and SMS providers, and CPU credits, because a $25 instance is burstable and only runs at 20% of each vCPU for free. Here is the whole bill for one AWS box keeping 30 days of logs, at AWS list prices (us-east-1, October 2026):

Per month10M100M518M (200/s, all month)
Server: t4g.medium, 2 vCPU / 4 GB$25$25$25
CPU credits above the burst baseline$0$0$31
Disk: gp3, 30 days of logs$3 (38 GB)$16 (199 GB)$76 (950 GB)
Backups: daily snapshots, kept a week$1$11$58
Bandwidth out: 10 KB a message, first 100 GB free$0$81$458
Total$29$133$648
Total sending through SES in the same region$29$52$190

Bandwidth is the big one at high volume. We assumed 10 KB leaves the box per notification, about one HTML email sent to a provider's API. SMS and push messages are a tenth of that. If your email provider is Amazon SES in the same region as the box, that row drops to zero, as long as the box has a public IP; from a private subnet, a NAT gateway charges $0.045 per GB unless you add a VPC endpoint for SES. Provider fees per message are extra in every case.

These rows are calculated, not measured. Disk uses the 1.9 KB per notification the database grew by during the 200/s soak, including open and click events. CPU credits use the 1.47 cores the box burned at 200/s.

What running it yourself takes

  • BackupsPostgres holds your users, templates and delivery log. Back it up like any production database.
  • Log pruningThe database grew by about 1.9 KB per notification, around 18 GB a month at 10 million. Keep as much history as you need and delete the rest.
  • Redis headroomRedis remembers each message for an hour so a retry is never sent twice. Size it for an hour of your peak traffic.
  • UpgradesNew versions ship as an npm package; bump it and redeploy. Database migrations run when the server starts.

When self-hosting makes sense

At low volume, hosted services are simpler and often free: Knock, Courier and Novu all include 10,000 messages a month. Usage-based pricing gets substantial at high volume, and that's where running it yourself starts to be worth the work.

Fig 4Monthly bill at list price

Hosted plans at their published per-message price, against one self-hosted box with everything on the bill from the table above. Provider fees for email, SMS and push are extra in every case.

  • Knock
  • Courier (dashed)
  • Novu Cloud
  • Self-hosted, one box
$0$100K$200K$300K$400K$500K0M20M40M60M80M100MKnock: $500,000Courier: $499,950Novu Cloud: $119,950Self-hosted: $133
Table view
Per monthKnock (Starter)Courier (Business)Novu Cloud (Team)Self-hosted
1M$5,000$4,950$1,150$26
10M$50,000$49,950$11,950$29
50M$250,000$249,950$59,950$75
100M$500,000$499,950$119,950$133

Prices are taken from each company's public pricing page, October 2026: Knock Starter, Courier Business, Novu Team, with the per-message overage extended to each volume. These are list prices. All three offer contracts at high volume, and a negotiated price can be much lower. Knock and Courier bill per message; Novu bills per workflow run. Knock and Courier differ by $50 at most, so their lines overlap.

Self-host notifkit if

  • You send enough that per-message pricing hurts.
  • Engineers own notifications. Templates and workflows are code in your repo, changed by pull request.
  • You already run Postgres and Redis. Or you don't mind adding two well-known services.
  • The data has to stay with you. Your users' details only leave your servers in the messages you send.

Use a hosted service if

  • Non-engineers edit the messages. Novu, Knock and Courier come with a visual editor; notifkit doesn't.
  • You want a drop-in in-app inbox. Notifkit can deliver to a webhook, but the UI is yours to build.
  • You don't want to run servers. Or be the one paged when they break.
  • Your volume is small. There's no bill to save.

Try it

If the left column sounds like you, the quickstart takes about ten minutes. You start a notifkit server next to Postgres and Redis, create a project, add a template and a user, and send your first notification. It's free and MIT-licensed, and your app can be written in any language: it only talks to notifkit over HTTP.

Send your first notification in ten minutes. Open the quickstart → GitHub

Reproduce it

The benchmark is in the notifkit repo under profiling/. Docker has to be running.

cd profiling

# 10-second tests: accept, catch up, keep up (3 runs per size)
A="--tiers=20,40,60 --ingest=10 --steady=10 --runs=3"
npm run profile:horizontal -- $A
npm run profile:vertical   -- $A

# Ramp from 150/s in +50/s one-minute steps, then a 15-minute soak at the highest that passed
npm run profile:sustained

# 15-minute soak at a fixed rate
npm run profile:sustained -- --soak-rate=200
Exact configuration

notifkit 1 vCPU / 1 GB running every service (API, enricher, engine, delivery, scheduler, events), Postgres 0.5 vCPU / 512 MB, Redis 0.5 vCPU with AOF (256 MB in the 10-second tests, 512 MB in the soaks). Worker concurrency 50, delivery concurrency 400, 5 Postgres connections. NODE_ENV=production, LOG_LEVEL=info. Load comes from a k6 container on the same Docker network. 100,000 seeded users, 8 templates, 35% critical, 40% normal and 25% low priority.

Will you get these numbers on a real server?

Not exactly, and we can't say which way. A server could be faster: no Windows VM layer, and the load generator isn't competing for the same CPU. It could be slower: a cloud vCPU is often slower per thread than a laptop core boosting past 4 GHz, burstable instances throttle, and real providers rate-limit you. The six 10-second runs of the $25 box varied by 5% (583 ± 28/s catching up), mostly from background load on the laptop. Each soak ran once.