Fourlab Insight · performance

The Bottleneck Is Usually Not Capacity

CSIRO’s Serverless Beacon is a useful reminder that performance problems are often workload problems in disguise. The real leverage is not always adding capacity. Sometimes it is making each query cheap enough that the business can ask more questions without carrying an always-on platform for every one of them.

2026-09-20

Photovisual Fourlab scene about The Bottleneck Is Usually Not Capacity: an operations surface with latency traces, customer-impact markers and one bottleneck made visible, with evidence cues for bottleneck, usually, latency.

The expensive part is often the wrong part

A lot of software leaders still talk about performance as if the answer is to add more of something: more nodes, more replicas, more compute, more headroom.

A current report about How CSIRO built scalable, cost-optimized genomic variant querying on AWS is the context here. The news is not the point; it makes the operational decision pressure visible.

That instinct is understandable. It is also often incomplete.

In practice, the bottleneck is frequently the shape of the workload. Which queries repeat. Which data is actually hot. Which parts of the stack are being kept alive because nobody has challenged the default architecture. Once a team starts paying for capacity before it understands the workload, the bill usually shows up twice: once in cloud spend, and again in operational drag.

A recent AWS architecture write-up about CSIRO’s Serverless Beacon makes that tradeoff visible in a concrete way. CSIRO, Australia’s national science agency, built a serverless solution for securely querying genomic variant data on AWS using Amazon S3, AWS Lambda, Amazon DynamoDB, and Amazon Athena. The point is not the domain itself. The point is the decision pressure: how do you support production-scale querying without turning the whole system into an always-on cost center?

That is the part many teams skip past too quickly.

Querying is a product decision, not just an infrastructure decision

The AWS example is about genomic variant querying, but the pattern is broader than life sciences.

A query layer is where product intent meets cost reality. Every search, lookup, and filter is a small tax on the system. If that tax is too high, people stop asking. They batch requests. They precompute too much. They avoid interactive workflows. The product gets slower in a way the dashboard does not always make obvious.

That is why CSIRO’s architecture is interesting. The team used serverless components to support production-scale clinical and research applications without building a traditional always-on platform first. That does not mean serverless is automatically the right answer for every workload. It does mean the architecture starts from a harder question: what do we actually need to keep warm, and what can stay cold until the business needs it?

For performance-minded leaders, that is a more useful question than “can it handle more traffic?”

Because traffic is not the same as bottleneck.

A system can absorb load and still be expensive to ask. It can look healthy on a dashboard and still be the reason teams hesitate to use it in the flow of work. In that sense, performance is not just a technical metric. It is a product constraint.

The scene most teams recognize

Picture a team review on a Tuesday morning.

The dashboard is fine. Latency is acceptable. Spend is creeping up. Someone proposes a larger cluster, a higher-tier database, or another round of caching.

Then somebody asks the annoying question: which queries are repeating?

That is usually where the room gets quiet.

Not because the team lacks talent. Because the answer is often uncomfortable. A small set of paths is carrying most of the load. A few expensive requests are being treated like a universal problem. Or the team is paying for broad capacity when only a narrow slice of the data is actually active.

This is where performance work becomes a leadership issue.

If you do not know which workload is dominant, you can build a faster system and still miss the business problem. You may improve throughput and leave the product expensive to operate. You may reduce latency and still keep the wrong thing always-on.

That is the hidden risk in a lot of “scale” conversations. The team thinks it is buying resilience, but it may actually be buying time to avoid a harder architectural decision.

What CSIRO’s architecture points to

The source is specific: sBeacon uses S3, Lambda, DynamoDB, and Athena to implement the GA4GH Beacon standard for secure genomic variant querying.

That stack says something important.

S3 is not there as a vanity choice. Lambda is not there because serverless is fashionable. DynamoDB is not there to impress an architecture review. Athena is not there to make the diagram look modern. Each piece is doing a job in a system that needs to answer queries without carrying the full fixed cost of a traditional platform.

That is the real move.

You reduce the cost of each query path so the business can afford to ask more questions, more often.

For software leaders, this matters because it changes the center of gravity in the architecture discussion. Scaling is not always about making the engine stronger. Sometimes it is about making each question cheaper.

That shift changes what gets built, what gets cached, what gets precomputed, and what gets left alone.

It also changes how teams talk about risk. Not in the dramatic sense, but in the practical one: if a system is expensive to query, people will work around it. They will export data. They will duplicate logic. They will create shadow paths because the official one is too costly or too slow for everyday use. The architecture then becomes more complex than the original problem it was meant to simplify.

The smallest useful intervention

The temptation in a situation like this is to redesign everything.

That is rarely the first move worth making.

The smallest useful intervention is usually to measure the shape of demand before expanding the shape of capacity. Which queries dominate. Which ones are interactive versus batch. Which datasets are hot. Which paths are being kept alive because they are convenient, not because they are necessary.

That evidence changes the conversation.

If most of the load sits on a narrow set of requests, a broad infrastructure upgrade may be the wrong fix. If the workload is bursty, serverless or partially serverless patterns may be a better fit than a permanently provisioned stack. If the business only needs a subset of the data to be immediately queryable, then keeping everything warm is a design choice, not a requirement.

CSIRO’s example is useful precisely because it makes that tradeoff concrete. The team did not appear to start from “how do we maximize infrastructure?” They started from “how do we support secure, production-scale querying in a way that fits the workload?” That is a more disciplined way to think about performance.

And it is usually cheaper than adding another layer of capacity first and asking questions later.

The leadership consequence

If you keep treating performance as a capacity problem, you will keep buying the most visible fix.

It feels responsible. It is easy to explain. It buys time.

But it also tends to hide the real bottleneck for another quarter.

The better question is less flattering and more useful: where in your stack are you paying for capacity before you have identified the real workload?

That question is uncomfortable because it usually lands on decisions already made. A database choice. A platform assumption. A habit of keeping systems warm because nobody wants to be the person who turns something off.

Still, that is where the leverage is.

CSIRO’s sBeacon is a reminder that the highest-value performance work is often not heroic tuning. It is making the system cheap enough to ask, so the business can learn faster without carrying a fixed cost for every possible query.

That is a different kind of scalability. And for most teams, it is the one that actually matters.