docs(architecture): add performance and scaling page compared to PHP-FPM
This commit is contained in:
92
docs/architecture/performance-and-scaling.md
Normal file
92
docs/architecture/performance-and-scaling.md
Normal file
@@ -0,0 +1,92 @@
|
||||
---
|
||||
title: Performance and scaling
|
||||
description: "Why a SummerCMS binary serves requests faster and with less memory than WinterCMS on PHP-FPM, where it does not, and how to scale and measure it."
|
||||
section: architecture
|
||||
order: 50
|
||||
---
|
||||
# Performance and scaling
|
||||
|
||||
WinterCMS runs under PHP-FPM, which boots the application again for every request. A SummerCMS application is one Go process that boots once and then serves every request from memory. This page explains where the speed and memory savings come from, where a port will not get faster, and how to scale and measure an application.
|
||||
|
||||
The figures on this page are typical for PHP-FPM with Laravel compared with Go services. SummerCMS has no published benchmarks yet, so treat them as expectations, not promises, and use [Measuring it yourself](#measuring-it-yourself) to get real numbers for your application.
|
||||
|
||||
## Boot once, not per request
|
||||
|
||||
PHP-FPM is shared-nothing: every request starts from an empty process state and re-runs the Laravel and Winter bootstrap. Service providers load, plugins register, configuration is read, routes are collected and the YAML caches are looked up. Even with OPcache this typically costs about 15-60 ms before your controller runs.
|
||||
|
||||
SummerCMS does that work once, at start-up. [party](../../modules/party/README.md) runs every plugin's Register and Boot (`party.Activate`), [cabana](../../modules/cabana/README.md) compiles each plugin's `fields.yaml` and `columns.yaml` into a `cabana.CompiledController`, and [surf](../../modules/surf/README.md) assembles every route into one standard library `http.ServeMux` (`surf.Assemble`). A request then only matches a route and runs its middleware chain, which costs microseconds. [Application lifecycle](application-lifecycle.md#start-up-sequence) lists the start-up steps and [Request lifecycle](request-lifecycle.md) follows a request through the router.
|
||||
|
||||
## Memory
|
||||
|
||||
An FPM worker typically holds 30-80 MB, and FPM needs one worker for every request it serves at the same time. Twenty workers therefore take roughly 1-1.5 GB before the database or the cache is counted. One SummerCMS process typically sits at 50-150 MB while it serves thousands of concurrent requests, because every request shares the same compiled routes, admin schemas and configuration.
|
||||
|
||||
## Slow I/O and concurrency
|
||||
|
||||
A PHP request that waits on an external API, an SMTP server, the search engine or the realtime server holds its FPM worker for the whole wait. When every worker is busy (`pm.max_children` is reached), new requests queue in front of FPM and p99 latency climbs sharply, even though the server's CPUs are idle.
|
||||
|
||||
In Go each request runs on its own goroutine, which costs a few KB of memory. A request that waits on the network parks its goroutine, and the runtime keeps serving other requests on the same threads, so slow upstreams do not block unrelated routes.
|
||||
|
||||
## Database connections
|
||||
|
||||
PHP-FPM usually opens one Postgres connection per worker, so the connection count grows with `pm.max_children` and with every server you add.
|
||||
|
||||
SummerCMS opens one `database/sql` pool per process. [lagoon](../../modules/lagoon/README.md) publishes it with `lagoon.Publish`, GORM queries run on it, and [conga](../../modules/conga/README.md) runs River jobs on the same pool. Connections are shared between requests and jobs instead of being held by idle workers. When the job worker runs, it adds one dedicated connection that listens for new jobs; [Running workers](../services/jobs.md#running-workers) explains it and what it needs from PgBouncer.
|
||||
|
||||
## Model hydration
|
||||
|
||||
Eloquent hydrates every row into a model object: it builds attribute arrays and runs magic accessors, mutators and casts for each model. GORM scans rows into plain Go structs through reflection, with no per-attribute magic, so each row costs less to load. The difference grows with the number of rows a route returns.
|
||||
|
||||
## What to expect
|
||||
|
||||
> [!NOTE]
|
||||
> The ranges below are typical for PHP-FPM with Laravel compared with Go services. They are not measurements of SummerCMS, which has no benchmarks yet. [Measuring it yourself](#measuring-it-yourself) shows how to get real numbers for your application.
|
||||
|
||||
| Endpoint shape | Examples | Typical gain over PHP-FPM |
|
||||
|----------------|----------|---------------------------|
|
||||
| Trivial or cached responses | Health checks, settings or lookup endpoints | 10-30x throughput per core, with sub-millisecond latency in Go |
|
||||
| Typical CRUD | Paginated lists, a record with a relation or two, validated writes | 3-10x |
|
||||
| Database-heavy | Large joins, aggregates, full-text queries | 1.2-2x, because time spent in Postgres dominates |
|
||||
|
||||
The less time a route spends waiting on Postgres, the more of the PHP bootstrap and hydration cost the port removes.
|
||||
|
||||
## Where you will not win
|
||||
|
||||
Some costs come over with the port unchanged, and a few framework behaviours need attention once you run more than one instance:
|
||||
|
||||
- A line-by-line port keeps the PHP queries, so slow queries and N+1 patterns come over unchanged. Fix them, for example with GORM's `Preload`, once the port passes parity.
|
||||
- [wire](../../modules/wire/README.md) keeps response bodies byte-compatible with the PHP backend, so payload sizes and the client's parsing work do not change.
|
||||
- Revoked tokens must be visible to every instance. Use `bouncer.NewPostgresBlacklist` (a `bouncer.PostgresBlacklist`) rather than the in-memory blacklist, so a logout on one replica applies on all of them; see [Refreshing and revoking](../services/authentication.md#refreshing-and-revoking).
|
||||
- `serve` runs the conga job worker in the same process as the HTTP server by default. Under load, set `queue.work_in_serve` to `false` and run `./bin/acme queue:work` as separate processes, so long jobs do not compete with requests for CPU; see [Running workers](../services/jobs.md#running-workers).
|
||||
- Thumbnails are generated in pure Go (`attach.File.Thumb`), which can be slower than GD or Imagick for large images.
|
||||
- Garbage collection pauses are sub-millisecond and do not matter at this scale.
|
||||
|
||||
> [!WARNING]
|
||||
> Rate limits are counted per process. The throttle middleware that `serve` builds keeps its counters in memory (`surf.MemoryStore`, behind the `surf.Store` interface), and `serve` offers no way to supply a different store yet. With N replicas the effective limit is N times the configured one. Until a shared store exists, divide the per-route limits by the replica count, or enforce the limit at the load balancer. [Rate limiting](../services/rate-limiting.md#named-buckets) describes the buckets.
|
||||
|
||||
## Scaling and operations
|
||||
|
||||
The application is one stateless binary. It starts in milliseconds, ships in a small container image and needs no OPcache warm-up after a deploy.
|
||||
|
||||
- Vertical scaling is more cores. The Go runtime uses them all, so there is no worker count to tune.
|
||||
- Horizontal scaling is N replicas behind a load balancer that share Postgres. Uploads must be shared too: point `storage.uploads.bucket_url` at a `file://` directory on storage that every replica mounts (see [Storage](../services/storage.md#bucket-urls)). Apply the blacklist and rate-limit notes above before you add the second replica.
|
||||
- Centrifugo scales on its own. [lighthouse](../../modules/lighthouse/README.md) only publishes to it and issues connection tokens, so adding application replicas does not change the realtime server; see [The Centrifugo driver](../services/realtime.md#the-centrifugo-driver).
|
||||
|
||||
A rollout runs the migrations once, then starts each replica behind the reverse proxy:
|
||||
|
||||
```sh
|
||||
./bin/acme migrate
|
||||
./bin/acme serve --addr 127.0.0.1:8080
|
||||
```
|
||||
|
||||
## Measuring it yourself
|
||||
|
||||
Compare the two backends on the routes your frontend actually calls. First record fixtures from the PHP backend with [tide](../../modules/tide/README.md) and replay them against the port, so you know both return the same responses:
|
||||
|
||||
```sh
|
||||
summer parity:record --spec testdata/parity/posts.spec.yaml --target http://127.0.0.1:8000 --output testdata/parity/posts.yaml --vars /tmp/parity/vars.yaml
|
||||
summer parity:replay --fixtures testdata/parity --target http://127.0.0.1:8080 --vars /tmp/parity/vars.yaml
|
||||
```
|
||||
|
||||
[The parity commands](../services/parity-testing.md#the-parity-commands) explains the flags and the variables file.
|
||||
|
||||
Once replay passes, drive the same routes under load against both backends with a load tool such as vegeta or k6. Run both on identical hardware against the same Postgres instance, with OPcache enabled and warm on the PHP side. For each route, compare p50 and p99 latency, requests per second and the resident memory (RSS) of the PHP-FPM pool and of the Go process.
|
||||
Reference in New Issue
Block a user