{"id":497799,"date":"2026-10-07T20:51:06","date_gmt":"2026-10-07T17:51:06","guid":{"rendered":"https:\/\/savepearlharbor.com\/?p=497799"},"modified":"2026-10-07T20:51:06","modified_gmt":"2026-10-07T17:51:06","slug":"","status":"publish","type":"post","link":"https:\/\/savepearlharbor.com\/?p=497799","title":{"rendered":"I Sent the Same PostgreSQL Traffic in Smooth and Bursty Patterns. Average Load Lied"},"content":{"rendered":"<div xmlns=\"http:\/\/www.w3.org\/1999\/xhtml\">\n<p>After experimenting with connection pool behavior, I wanted to test another assumption that looks harmless on dashboards.<\/p>\n<p>If two workloads have the same average request rate, are they really equivalent from the database point of view?<\/p>\n<p>At first, it seems reasonable.<\/p>\n<p>If one service receives 1,000 requests per second on average and another also receives 1,000 requests per second, both appear to place roughly the same load on PostgreSQL.<\/p>\n<p>But averages remove timing.<\/p>\n<p>And timing is exactly where queues appear.<\/p>\n<p>So I built a controlled queueing simulation of a Go service talking to PostgreSQL through a connection pool.<\/p>\n<p>This was not a benchmark of PostgreSQL itself. I deliberately kept the model simple because I wanted to isolate one variable: the arrival pattern.<\/p>\n<p>The total amount of traffic stayed the same.<\/p>\n<p>Only the timing changed.<\/p>\n<p>One workload was smooth.<\/p>\n<p>The other arrived in bursts.<\/p>\n<p>The average RPS was identical.<\/p>\n<p>The model behaved very differently.<\/p>\n<figure class=\"full-width \"><img decoding=\"async\" src=\"https:\/\/habrastorage.org\/r\/w1560\/getpro\/habr\/upload_files\/77e\/0e4\/600\/77e0e46009f1bd2aacab647528bcc42e.png\" width=\"1672\" height=\"941\" sizes=\"auto, (max-width: 780px) 100vw, 50vw\" srcset=\"https:\/\/habrastorage.org\/r\/w780\/getpro\/habr\/upload_files\/77e\/0e4\/600\/77e0e46009f1bd2aacab647528bcc42e.png 780w,&#10;       https:\/\/habrastorage.org\/r\/w1560\/getpro\/habr\/upload_files\/77e\/0e4\/600\/77e0e46009f1bd2aacab647528bcc42e.png 781w\" loading=\"lazy\" decode=\"async\"\/><\/figure>\n<h3>What I wanted to isolate<\/h3>\n<p>A real PostgreSQL system has too many moving parts for this particular question.<\/p>\n<p>Query plans change.<\/p>\n<p>Caches warm up.<\/p>\n<p>Checkpoints happen.<\/p>\n<p>Autovacuum runs.<\/p>\n<p>Storage latency varies.<\/p>\n<p>CPU scheduling changes.<\/p>\n<p>Background processes compete for resources.<\/p>\n<p>If I ran the same experiment directly against a production-like database, a change in p99 would not automatically prove that traffic burstiness caused it.<\/p>\n<p>So I started with a narrower model.<\/p>\n<p>The simulated application received approximately 1,000 requests per second on average.<\/p>\n<p>Every request required one database connection.<\/p>\n<p>Most operations represented short indexed queries.<\/p>\n<p>A smaller group represented slower operations.<\/p>\n<p>The query-time distribution stayed unchanged between runs.<\/p>\n<p>The connection pool size stayed unchanged.<\/p>\n<p>The test duration stayed unchanged.<\/p>\n<p>The total number of requests stayed unchanged.<\/p>\n<p>Only their arrival times changed.<\/p>\n<p>I tracked:<\/p>\n<ul>\n<li>\n<p>incoming RPS<\/p>\n<\/li>\n<li>\n<p>completed RPS<\/p>\n<\/li>\n<li>\n<p>connection acquisition time<\/p>\n<\/li>\n<li>\n<p>active connections<\/p>\n<\/li>\n<li>\n<p>request p50<\/p>\n<\/li>\n<li>\n<p>request p95<\/p>\n<\/li>\n<li>\n<p>request p99<\/p>\n<\/li>\n<li>\n<p>queue depth over time<\/p>\n<\/li>\n<\/ul>\n<p>The exact latency values are not intended as PostgreSQL performance numbers.<\/p>\n<p>The useful part of the simulation is the behavior of the queue when the same amount of work arrives in a different temporal pattern.<\/p>\n<h3>The smooth workload was boring<\/h3>\n<p>That was actually a good sign.<\/p>\n<p>With traffic distributed relatively evenly, the connection pool rarely experienced sustained pressure.<\/p>\n<p>The number of active connections moved around, but the queue remained small.<\/p>\n<p>Temporary waiting appeared and disappeared quickly.<\/p>\n<p>Completed throughput closely followed incoming traffic.<\/p>\n<p>Latency remained stable.<\/p>\n<p>The important part was not the exact number of milliseconds.<\/p>\n<p>It was the shape.<\/p>\n<p>Work arrived.<\/p>\n<p>The system processed it.<\/p>\n<p>Nothing accumulated for long.<\/p>\n<p>In queueing terms, the model remained stable.<\/p>\n<h3>Then I changed only the arrival pattern<\/h3>\n<p>For the second run, I kept the same total number of requests but grouped them into short spikes.<\/p>\n<p>Instead of approximately 1,000 requests arriving every second, the traffic could look more like this over several consecutive intervals:<\/p>\n<p>1,700<\/p>\n<p>1,600<\/p>\n<p>400<\/p>\n<p>300<\/p>\n<p>1,100<\/p>\n<p>900<\/p>\n<p>Over a longer window, the average remained close to 1,000 RPS.<\/p>\n<p>The amount of work was essentially unchanged.<\/p>\n<p>The timing was not.<\/p>\n<p>During each spike, requests arrived faster than the model could immediately process them.<\/p>\n<p>A queue formed.<\/p>\n<p>When traffic dropped, the system started draining it.<\/p>\n<p>Sometimes the queue disappeared before the next spike.<\/p>\n<p>Sometimes the next burst arrived while part of the previous queue still existed.<\/p>\n<p>That was enough to change tail latency significantly.<\/p>\n<h3>Average RPS hid almost everything<\/h3>\n<p>Across the complete run, both workloads had nearly the same average incoming rate.<\/p>\n<p>A dashboard showing one-minute averages could make them look almost identical.<\/p>\n<p>That was the point of the experiment.<\/p>\n<p>Average throughput told me how much work arrived during the window.<\/p>\n<p>It did not tell me when that work arrived.<\/p>\n<p>The smooth workload used available service capacity gradually.<\/p>\n<p>The bursty workload repeatedly exceeded short-term capacity and created waiting.<\/p>\n<p>The total work could be the same while the user-visible behavior was very different.<\/p>\n<h3>p50 barely moved<\/h3>\n<p>Median latency was not the first metric to react.<\/p>\n<p>Most requests still experienced little or no queueing.<\/p>\n<p>During quiet periods, a request could enter immediately.<\/p>\n<p>At the beginning of a burst, some requests also got a connection before the pool filled.<\/p>\n<p>So p50 remained relatively stable.<\/p>\n<p>That created a combination I find dangerous in real monitoring:<\/p>\n<p>average RPS looks normal<\/p>\n<p>median latency looks normal<\/p>\n<p>most individual operations still look fast<\/p>\n<p>and a small percentage of requests experience much worse latency<\/p>\n<p>The problem appeared much more clearly in p95 and p99.<\/p>\n<h3>p99 reacted to queue position<\/h3>\n<p>Under smooth traffic, requests saw relatively similar conditions.<\/p>\n<p>Under bursty traffic, arrival timing became important.<\/p>\n<p>Imagine two requests with identical simulated query times.<\/p>\n<p>The first arrives while a connection is available.<\/p>\n<p>The second arrives a few milliseconds later, after the pool has become saturated.<\/p>\n<p>The first begins immediately.<\/p>\n<p>The second waits behind other work.<\/p>\n<p>The database operation itself has not changed.<\/p>\n<p>End-to-end latency has.<\/p>\n<p>That is why tail latency became much more sensitive to burstiness than median latency.<\/p>\n<p>The model did not need slower queries to produce slower requests.<\/p>\n<p>Queueing was enough.<\/p>\n<h3>Same amount of work, different queue<\/h3>\n<p>This became the simplest description of the result.<\/p>\n<p>The workload performed roughly the same total amount of work.<\/p>\n<p>The difference was how that work was presented to the system.<\/p>\n<p>With smooth arrivals, capacity was consumed gradually.<\/p>\n<p>With bursts, periods of spare capacity could not compensate for later spikes.<\/p>\n<p>If the system has unused capacity during one quiet second, it cannot store that capacity and use it during the next overloaded second.<\/p>\n<p>That unused capacity disappears.<\/p>\n<p>Requests arriving later still have to wait.<\/p>\n<h3>Capacity is not a bank account<\/h3>\n<p>This is where average utilization can become misleading.<\/p>\n<p>Suppose a service can sustainably process around 1,100 requests per second.<\/p>\n<p>Its average traffic is 800 RPS.<\/p>\n<p>It is tempting to describe that as roughly 300 RPS of spare capacity.<\/p>\n<p>That can be useful for rough planning.<\/p>\n<p>But it says nothing about short spikes.<\/p>\n<p>If traffic temporarily reaches 1,500 RPS, the unused capacity from the previous quiet period does not help.<\/p>\n<p>The system has instantaneous capacity.<\/p>\n<p>Not stored capacity.<\/p>\n<p>That distinction is especially important for tail latency.<\/p>\n<p>A service can be comfortably below capacity on average and still create repeated queues.<\/p>\n<h3>The connection pool made the mechanism visible<\/h3>\n<p>In the simulation, the connection pool acted as the admission boundary.<\/p>\n<p>Before all slots were occupied, requests could proceed immediately.<\/p>\n<p>Once the pool became saturated, new requests waited.<\/p>\n<p>This produced two clearly different components of latency:<\/p>\n<p>connection acquisition time<\/p>\n<p>and execution time after acquisition<\/p>\n<p>That separation matters.<\/p>\n<p>Execution time answers one question:<\/p>\n<p>How long did the work take after it started?<\/p>\n<p>Acquisition time answers another:<\/p>\n<p>How much work was already ahead of this request?<\/p>\n<p>A request can therefore become slow even when the simulated database operation remains fast.<\/p>\n<h3>A larger pool would not automatically solve the experiment<\/h3>\n<p>One obvious response to bursty traffic is to increase the pool size.<\/p>\n<p>That might reduce application-side waiting.<\/p>\n<p>But this simulation deliberately does not model PostgreSQL contention, so it cannot prove whether a larger pool would improve or worsen a real database workload.<\/p>\n<p>What it can show is simpler.<\/p>\n<p>Increasing the pool changes where the concurrency limit exists.<\/p>\n<p>If the database has enough unused capacity, allowing more concurrent work may help.<\/p>\n<p>If the database is already close to its efficient concurrency limit, a larger pool may only move the queue deeper into the system.<\/p>\n<p>That second part needs a real PostgreSQL benchmark to verify.<\/p>\n<p>I would not infer it from this simulation alone.<\/p>\n<h3>Monitoring resolution matters<\/h3>\n<p>The model also made aggregation effects easy to see.<\/p>\n<p>Suppose traffic over several short intervals is:<\/p>\n<p>500 RPS<\/p>\n<p>600 RPS<\/p>\n<p>1,800 RPS<\/p>\n<p>1,900 RPS<\/p>\n<p>500 RPS<\/p>\n<p>700 RPS<\/p>\n<p>A wide enough time window can compress this into a completely ordinary average.<\/p>\n<p>The spikes disappear visually.<\/p>\n<p>The queue does not.<\/p>\n<p>The same principle applies to many production metrics.<\/p>\n<p>CPU can briefly saturate and still show a moderate one-minute average.<\/p>\n<p>Active connections can repeatedly hit the limit while a longer aggregation window looks stable.<\/p>\n<p>p99 can spike for short periods without dramatically moving a long-term average.<\/p>\n<p>For burst-sensitive workloads, I would want short-interval metrics placed on the same timeline.<\/p>\n<p>Especially:<\/p>\n<ul>\n<li>\n<p>incoming RPS<\/p>\n<\/li>\n<li>\n<p>completed RPS<\/p>\n<\/li>\n<li>\n<p>active connections<\/p>\n<\/li>\n<li>\n<p>connection acquisition latency<\/p>\n<\/li>\n<li>\n<p>queue depth<\/p>\n<\/li>\n<li>\n<p>request p95 and p99<\/p>\n<\/li>\n<\/ul>\n<p>The relationship between those metrics matters more than any single number.<\/p>\n<h3>Bursts do not have to come from users<\/h3>\n<p>Another useful implication is that uneven database traffic can be created internally.<\/p>\n<p>Examples include:<\/p>\n<p>scheduled jobs<\/p>\n<p>batch workers<\/p>\n<p>cache expiration<\/p>\n<p>message consumers restarting<\/p>\n<p>periodic synchronization<\/p>\n<p>retry waves<\/p>\n<p>multiple workers waking up on the same timer<\/p>\n<p>A system can therefore have stable external traffic and still generate highly uneven database demand.<\/p>\n<p>Imagine 100 workers performing the same periodic task every minute.<\/p>\n<p>Their average database traffic can be tiny.<\/p>\n<p>But if all 100 wake up at the same second, instantaneous concurrency can be large.<\/p>\n<p>The average does not describe that behavior well.<\/p>\n<h3>Smoothing work can be an optimization<\/h3>\n<p>This is probably the most practical conclusion I would take from the model.<\/p>\n<p>Performance optimization does not always mean doing less work or making each operation faster.<\/p>\n<p>Sometimes it means changing when the same work happens.<\/p>\n<p>Imagine a batch job that needs to perform 10,000 database operations.<\/p>\n<p>One approach is to launch as much concurrency as possible and finish quickly.<\/p>\n<p>Another is to use a bounded worker pool and spread the operations over a longer interval.<\/p>\n<p>The amount of useful work is identical.<\/p>\n<p>The second version may produce much less interference with latency-sensitive requests.<\/p>\n<p>The batch takes longer.<\/p>\n<p>The overall service can behave better.<\/p>\n<p>That is load shaping rather than load reduction.<\/p>\n<h3>Backpressure matters more under bursty traffic<\/h3>\n<p>Bursts are especially useful for exposing systems without effective backpressure.<\/p>\n<p>When the arrival rate exceeds service capacity, something has to happen.<\/p>\n<p>Requests can queue.<\/p>\n<p>They can be rejected.<\/p>\n<p>Producers can slow down.<\/p>\n<p>Workers can limit concurrency.<\/p>\n<p>Or the system can continue accepting unlimited work until another resource becomes the bottleneck.<\/p>\n<p>The first four options are controlled behaviors.<\/p>\n<p>The last one usually is not.<\/p>\n<p>A connection pool is one form of backpressure.<\/p>\n<p>A bounded queue is another.<\/p>\n<p>A rate limiter is another.<\/p>\n<p>The correct mechanism depends on the system, but the goal is similar:<\/p>\n<p>prevent short demand spikes from turning into unbounded concurrency.<\/p>\n<h3>Retries could make the effect much worse<\/h3>\n<p>I intentionally excluded retries from this simulation.<\/p>\n<p>They would make it harder to isolate the original arrival pattern.<\/p>\n<p>But they are an obvious next step.<\/p>\n<p>Suppose a burst increases queueing enough that some requests time out.<\/p>\n<p>Clients retry.<\/p>\n<p>Now additional traffic appears specifically because the system is already under pressure.<\/p>\n<p>Latency causes retries.<\/p>\n<p>Retries cause more load.<\/p>\n<p>More load causes more latency.<\/p>\n<p>If many clients use the same timeout and retry interval, the retries themselves can become synchronized.<\/p>\n<p>That can turn a short external spike into a longer internally generated wave.<\/p>\n<p>Adding exponential backoff and jitter is one way to reduce that synchronization.<\/p>\n<p>Again, the total amount of useful work has not necessarily changed.<\/p>\n<p>The timing has.<\/p>\n<h3>Average RPS is not a workload description<\/h3>\n<p>This was the main thing the experiment changed for me.<\/p>\n<p>I used to think about RPS as one of the primary descriptions of workload size.<\/p>\n<p>Now I think average RPS is only one property.<\/p>\n<p>Two services can both average 1,000 requests per second and have completely different operational characteristics.<\/p>\n<p>One might remain between 900 and 1,100 almost continuously.<\/p>\n<p>Another might alternate between 200 and 1,800.<\/p>\n<p>Their averages are similar.<\/p>\n<p>Their queueing behavior is not.<\/p>\n<p>Their required headroom is not.<\/p>\n<p>Their p99 behavior is not.<\/p>\n<p>Their timeout sensitivity is not.<\/p>\n<p>Their connection pool behavior is not.<\/p>\n<p>That is a lot of information for one average to hide.<\/p>\n<h3>What I would test against real PostgreSQL<\/h3>\n<p>The simulation answers only the queueing question.<\/p>\n<p>The next step would be a real PostgreSQL benchmark.<\/p>\n<p>I would keep:<\/p>\n<p>the query mix<\/p>\n<p>the connection pool<\/p>\n<p>the total request count<\/p>\n<p>and the average request rate<\/p>\n<p>as stable as possible.<\/p>\n<p>Then I would run at least two arrival patterns:<\/p>\n<p>smooth<\/p>\n<p>and bursty<\/p>\n<p>On the Go side, I would collect:<\/p>\n<ul>\n<li>\n<p>incoming requests<\/p>\n<\/li>\n<li>\n<p>completed requests<\/p>\n<\/li>\n<li>\n<p>connection acquisition duration<\/p>\n<\/li>\n<li>\n<p>active connections<\/p>\n<\/li>\n<li>\n<p>idle connections<\/p>\n<\/li>\n<li>\n<p>pool waiters<\/p>\n<\/li>\n<li>\n<p>request p50<\/p>\n<\/li>\n<li>\n<p>request p95<\/p>\n<\/li>\n<li>\n<p>request p99<\/p>\n<\/li>\n<\/ul>\n<p>On PostgreSQL, I would collect:<\/p>\n<ul>\n<li>\n<p>active sessions<\/p>\n<\/li>\n<li>\n<p>query latency<\/p>\n<\/li>\n<li>\n<p>transactions per second<\/p>\n<\/li>\n<li>\n<p>CPU utilization<\/p>\n<\/li>\n<li>\n<p>I\/O latency<\/p>\n<\/li>\n<li>\n<p>lock waits<\/p>\n<\/li>\n<li>\n<p>relevant query statistics<\/p>\n<\/li>\n<\/ul>\n<p>That would answer a second question that this model cannot:<\/p>\n<p>Does burstiness only create application-side queueing, or does additional database concurrency also change query execution itself?<\/p>\n<p>That distinction matters.<\/p>\n<p>If acquisition latency rises while query execution remains stable, most of the problem is before execution.<\/p>\n<p>If query duration also rises during bursts, PostgreSQL itself is experiencing additional contention.<\/p>\n<h3>What I would try before adding more capacity<\/h3>\n<p>If production measurements showed that burstiness was the real cause, I would first ask whether the spikes could be reduced.<\/p>\n<p>Possible approaches include:<\/p>\n<p>adding jitter to scheduled jobs<\/p>\n<p>limiting background worker concurrency<\/p>\n<p>breaking large batches into smaller units<\/p>\n<p>using queues<\/p>\n<p>avoiding synchronized cache expiration<\/p>\n<p>rate limiting non-interactive work<\/p>\n<p>using exponential backoff for retries<\/p>\n<p>separating latency-sensitive and background workloads<\/p>\n<p>These changes do not necessarily reduce total work.<\/p>\n<p>They reduce simultaneous work.<\/p>\n<p>That distinction is important.<\/p>\n<h3>The next experiment<\/h3>\n<p>There is an obvious continuation.<\/p>\n<p>Keep the same bursty arrival pattern.<\/p>\n<p>Keep the same total work.<\/p>\n<p>Change only the admission policy.<\/p>\n<p>I would compare:<\/p>\n<p>one shared connection pool<\/p>\n<p>a smaller bounded pool<\/p>\n<p>separate pools for interactive and background traffic<\/p>\n<p>The interesting metric would not be maximum throughput.<\/p>\n<p>It would be whether interactive p99 changes while total completed work remains approximately the same.<\/p>\n<p>That would test whether connection pools can act not only as reusable connection storage, but also as a simple traffic scheduling mechanism.<\/p>\n<h3>What I took away from this<\/h3>\n<p>This simulation did not prove that bursty traffic makes PostgreSQL itself slower.<\/p>\n<p>It was not designed to.<\/p>\n<p>What it did show is that two workloads with the same average request rate can create very different queueing behavior before database execution even begins.<\/p>\n<p>The total amount of work stayed approximately the same.<\/p>\n<p>The average RPS stayed approximately the same.<\/p>\n<p>The timing changed.<\/p>\n<p>That alone was enough to change tail latency.<\/p>\n<p>So average RPS tells me how much work arrived during a window.<\/p>\n<p>It does not tell me how that work arrived.<\/p>\n<p>And when a database-backed service operates close to a concurrency limit, that missing information can be more important than the average itself.<\/p>\n<\/div>\n<p>\u0441\u0441\u044b\u043b\u043a\u0430 \u043d\u0430 \u043e\u0440\u0438\u0433\u0438\u043d\u0430\u043b \u0441\u0442\u0430\u0442\u044c\u0438 <a href=\"https:\/\/habr.com\/ru\/articles\/1091752\/\">https:\/\/habr.com\/ru\/articles\/1091752\/<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>After experimenting with connection pool behavior, I wanted to test another assumption that looks harmless on dashboards.If two workloads have the same average request rate, are they really equivalent from the database point of view?At first, it seems reasonable.If one service receives 1,000 requests per second on average and another also receives 1,000 requests per second, both appear to place roughly the same load on PostgreSQL.But averages remove timing.And timing is exactly where queues appear.So I built a controlled queueing simulation of a Go service talking to PostgreSQL through a connection pool.This was not a benchmark of PostgreSQL itself. I deliberately kept the model simple because I wanted to isolate one variable: the arrival pattern.The total amount of traffic stayed the same.Only the timing changed.One workload was smooth.The other arrived in bursts.The average RPS was identical.The model behaved very differently.What I wanted to isolateA real PostgreSQL system has too many moving parts for this particular question.Query plans change.Caches warm up.Checkpoints happen.Autovacuum runs.Storage latency varies.CPU scheduling changes.Background processes compete for resources.If I ran the same experiment directly against a production-like database, a change in p99 would not automatically prove that traffic burstiness caused it.So I started with a narrower model.The simulated application received approximately 1,000 requests per second on average.Every request required one database connection.Most operations represented short indexed queries.A smaller group represented slower operations.The query-time distribution stayed unchanged between runs.The connection pool size stayed unchanged.The test duration stayed unchanged.The total number of requests stayed unchanged.Only their arrival times changed.I tracked:incoming RPScompleted RPSconnection acquisition timeactive connectionsrequest p50request p95request p99queue depth over timeThe exact latency values are not intended as PostgreSQL performance numbers.The useful part of the simulation is the behavior of the queue when the same amount of work arrives in a different temporal pattern.The smooth workload was boringThat was actually a good sign.With traffic distributed relatively evenly, the connection pool rarely experienced sustained pressure.The number of active connections moved around, but the queue remained small.Temporary waiting appeared and disappeared quickly.Completed throughput closely followed incoming traffic.Latency remained stable.The important part was not the exact number of milliseconds.It was the shape.Work arrived.The system processed it.Nothing accumulated for long.In queueing terms, the model remained stable.Then I changed only the arrival patternFor the second run, I kept the same total number of requests but grouped them into short spikes.Instead of approximately 1,000 requests arriving every second, the traffic could look more like this over several consecutive intervals:1,7001,6004003001,100900Over a longer window, the average remained close to 1,000 RPS.The amount of work was essentially unchanged.The timing was not.During each spike, requests arrived faster than the model could immediately process them.A queue formed.When traffic dropped, the system started draining it.Sometimes the queue disappeared before the next spike.Sometimes the next burst arrived while part of the previous queue still existed.That was enough to change tail latency significantly.Average RPS hid almost everythingAcross the complete run, both workloads had nearly the same average incoming rate.A dashboard showing one-minute averages could make them look almost identical.That was the point of the experiment.Average throughput told me how much work arrived during the window.It did not tell me when that work arrived.The smooth workload used available service capacity gradually.The bursty workload repeatedly exceeded short-term capacity and created waiting.The total work could be the same while the user-visible behavior was very different.p50 barely movedMedian latency was not the first metric to react.Most requests still experienced little or no queueing.During quiet periods, a request could enter immediately.At the beginning of a burst, some requests also got a connection before the pool filled.So p50 remained relatively stable.That created a combination I find dangerous in real monitoring:average RPS looks normalmedian latency looks normalmost individual operations still look fastand a small percentage of requests experience much worse latencyThe problem appeared much more clearly in p95 and p99.p99 reacted to queue positionUnder smooth traffic, requests saw relatively similar conditions.Under bursty traffic, arrival timing became important.Imagine two requests with identical simulated query times.The first arrives while a connection is available.The second arrives a few milliseconds later, after the pool has become saturated.The first begins immediately.The second waits behind other work.The database operation itself has not changed.End-to-end latency has.That is why tail latency became much more sensitive to burstiness than median latency.The model did not need slower queries to produce slower requests.Queueing was enough.Same amount of work, different queueThis became the simplest description of the result.The workload performed roughly the same total amount of work.The difference was how that work was presented to the system.With smooth arrivals, capacity was consumed gradually.With bursts, periods of spare capacity could not compensate for later spikes.If the system has unused capacity during one quiet second, it cannot store that capacity and use it during the next overloaded second.That unused capacity disappears.Requests arriving later still have to wait.Capacity is not a bank accountThis is where average utilization can become misleading.Suppose a service can sustainably process around 1,100 requests per second.Its average traffic is 800 RPS.It is tempting to describe that as roughly 300 RPS of spare capacity.That can be useful for rough planning.But it says nothing about short spikes.If traffic temporarily reaches 1,500 RPS, the unused capacity from the previous quiet period does not help.The system has instantaneous capacity.Not stored capacity.That distinction is especially important for tail latency.A service can be comfortably below capacity on average and still create repeated queues.The connection pool made the mechanism visibleIn the simulation, the connection pool acted as the admission boundary.Before all slots were occupied, requests could proceed immediately.Once the pool became saturated, new requests waited.This produced two clearly different components of latency:connection acquisition timeand execution time after acquisitionThat separation matters.Execution time answers one question:How long did the work take after it started?Acquisition time answers another:How much work was already ahead of this request?A request can therefore become slow even when the simulated database operation remains fast.A larger pool would not automatically solve the experimentOne obvious response to bursty traffic is to increase the pool size.That might reduce application-side waiting.But this simulation deliberately does not model PostgreSQL contention, so it cannot prove whether a larger pool would improve or worsen a real database workload.What it can show is simpler.Increasing the pool changes where the concurrency limit exists.If the database has enough unused capacity, allowing more concurrent work may help.If the database is already close to its efficient concurrency limit, a larger pool may only move the queue deeper into the system.That second part needs a real PostgreSQL benchmark to verify.I would not infer it from this simulation alone.Monitoring resolution mattersThe model also made aggregation effects easy to see.Suppose traffic over several short intervals is:500 RPS600 RPS1,800 RPS1,900 RPS500 RPS700 RPSA wide enough time window can compress this into a completely ordinary average.The spikes disappear visually.The queue does not.The same principle applies to many production metrics.CPU can briefly saturate and still show a moderate one-minute average.Active connections can repeatedly hit the limit while a longer aggregation window looks stable.p99 can spike for short periods without dramatically moving a long-term average.For burst-sensitive workloads, I would want short-interval metrics placed on the same timeline.Especially:incoming RPScompleted RPSactive connectionsconnection acquisition latencyqueue depthrequest p95 and p99The relationship between those metrics matters more than any single number.Bursts do not have to come from usersAnother useful implication is that uneven database traffic can be created internally.Examples include:scheduled jobsbatch workerscache expirationmessage consumers restartingperiodic synchronizationretry wavesmultiple workers waking up on the same timerA system can therefore have stable external traffic and still generate highly uneven database demand.Imagine 100 workers performing the same periodic task every minute.Their average database traffic can be tiny.But if all 100 wake up at the same second, instantaneous concurrency can be large.The average does not describe that behavior well.Smoothing work can be an optimizationThis is probably the most practical conclusion I would take from the model.Performance optimization does not always mean doing less work or making each operation faster.Sometimes it means changing when the same work happens.Imagine a batch job that needs to perform 10,000 database operations.One approach is to launch as much concurrency as possible and finish quickly.Another is to use a bounded worker pool and spread the operations over a longer interval.The amount of useful work is identical.The second version may produce much less interference with latency-sensitive requests.The batch&#8230;<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[],"tags":[],"class_list":["post-497799","post","type-post","status-publish","format-standard","hentry"],"_links":{"self":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts\/497799","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=497799"}],"version-history":[{"count":0,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts\/497799\/revisions"}],"wp:attachment":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=497799"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=497799"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=497799"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}