Scaling Cloud IDE Provisioning by 5x: From 5K to 25K Workspaces in One Region
February 21, 2026

HackerRank puts candidates in a real VS Code workspace, right in the browser. For a long time, one region could hold about 5,000 of those at once, and that was fine.
Then the questions changed. Hiring moved toward multi-file project work, the kind of setup people now use with coding agents: several files open, a real repo, a language server running, not one lonely editor pane. Demand for concurrent workspaces climbed fast. And 5,000 went from a comfortable ceiling to a hard one.
Past that point the platform just couldn’t keep up. Some of it was infrastructure: quotas, IP space, NFS. Some of it was our own code: VMs created one at a time, chatty APIs, a database that wasn’t ready for the write rate. I owned the work end to end, taking one region from 5,000 to 25,000 concurrent workspaces.
This was never a “spin up more VMs” problem.
The whole path was the ceiling
At this scale, provisioning is a coordination problem. One workspace request touches the cloud control plane, VM create, our database, git, NFS, and the service that keeps all of that in sync. If any one of those is slow or failing, the candidate sits staring at a spinner.

I didn’t get to pick a single villain. At 5,000, every one of those was already red. Fixing one service on its own didn’t move the overall ceiling at all. So the work turned into a loop: find where the system was actually breaking, fix that, load test again, watch the next thing catch fire.
GCP was the primary capacity. AWS was the fallback. That dual-cloud setup mattered later, when a bulk trick on one provider had no equivalent on the other.
First, the cloud had to stop saying no
Before I touched a line of service code, I needed headroom. Otherwise every load test would die on provider throttling, and I’d never get to see our own bugs.
Quota and capacity changes:
- Read requests per region/min: 15,000
- Write requests per region/min: 16,500
- Queries per region/min: 15,000
- Instance capacity per VPC: 50,000
- Private IP allocation: 65,000, by expanding the primary range to a
/16
That cleared the first hard blocker: the provider saying no before our code even had a chance to fail.
Bulk APIs, or you lose on round trips
Single-resource operations were too expensive at this concurrency. Each create, each route insert, each status fetch was a round trip, and we couldn’t afford those when thousands of workspaces were coming up at once.
VM creation moved to bulk insert APIs, which cut control-plane overhead on each cloud.
After create, route entries went in through Redis pipelines, in batches, instead of one round trip per route.
Reconciliation was the awkward part. There’s no bulk-read primitive that behaves the same on GCP and AWS, so I used filtered list APIs in controlled batches to keep our state lined up with what actually existed.
On the non-bulk path, fetching instance details used to take 3 API calls. I cut that to 1. On its own, that looks like a nothing-change. But during a burst, with retries piling on, that path got about 3x faster, and it stopped the slow path from amplifying the incident.
Then the database became the thing we were waiting on
Once VMs were coming up fast enough, the database showed up as the new ceiling. Connections ran out even while the cloud was keeping up. That was the first surprise: compute was no longer the scarce resource. The write path was.
I tightened connection pooling in the workspace service, so we weren’t opening and closing a connection on every state write.
Then I replaced one-by-one updates with bulk queries for workspace state and runtime fields. That dropped both the connection count and the write overhead.
A workspace is not provisioned until the repo is there
A VM with no project on it is still a failed workspace. Git repos for those IDEs lived on NFS, and that clone is what the candidate actually opens. That path had to take the same 5x, and NFS turned out to be the second surprise hiding behind VM create.
What changed:
- Right-sized node pools and replicas for the git service
- One atomic push into git instead of multiple round trips
- git repack configuration to cut NFS IOPS
- Load-tested NFS mount options for metadata pressure
- Regional NFS in production, with a migration plan behind it
Fast VMs with a slow clone still look like a provisioning failure to the candidate. NFS metadata IOPS is how that failure shows up: not a dramatic outage, just thousands of small file operations stepping on each other while the workspace is supposedly “ready.”
I watched the load test
I ran JMeter against production-like conditions. The target was 25,000 workspaces in one region, at about 1,000 assignments per minute. This wasn’t a pass/fail checkbox at the end. Each run told me what to fix next: hidden coupling, retry behavior, the NFS and database limits that only show up once the rest of the path is healthy.
The last run is this dashboard. I was watching it live.

What I was looking at:
- Running workspaces climbing through 25K
- Create requests holding around 1,000 per minute
- Average time to assign under a second (0.97s on this run)
- Two requests waiting, not a queue growing without bound
- 26,037 backends up
- 4 healthy contexts: the cloud-provider connections the workspace service can provision through. If that number drops, a provider is out of the pool and we’ve lost our fallback.
What this left me with
Scale here is a systems problem. The VM was never the whole story.
Bulk primitives become mandatory past a certain rate. Keep doing one-at-a-time against the cloud, the database, and git, and you’ll hit 5,000 again under a new name.
The long tail matters more than the architecture diagram suggests. The 3-to-1 API change, the NFS mount options, the bulk SQL: none of those show up when you load test at 500 workspaces.
If you don’t test like production, production will test you.