Synthetic Industry

Job backend-connection-pool-exhaustion-repaired · revised 11 October 2026

Stop "out of database connections" errors in one service, pool budgeted and tested

One service that times out waiting for a database connection, or hits the server limit, is diagnosed and corrected, then completes a load test at the concurrency you name.

You might be seeing

  • Errors say the pool timed out waiting for a connection, or the server has too many connections
  • Failures cluster at busy times and the service recovers after a restart
  • The database shows many connections that are idle or idle in a transaction

No passwords, keys, card details or admin invites needed to start.

What usually happened

Several different faults produce the same error. The pool may be smaller than the concurrency the service allows, connections may be taken and never returned, each process or worker may open its own pool so the total passes the server's limit, a pooler mode may not suit the features the code uses, or a request may hold a connection while it waits on something slow. Raising a limit hides the cause and can push the database over its own ceiling.

Who it’s for: An engineering lead whose service logs errors such as a timeout waiting for a connection or "too many connections" when traffic rises.

Usually starts when: Requests fail in bursts with connection errors, restarting the service clears it for a while, and raising a limit has only moved the problem.

The result: Driven by a load test at the concurrency you name, the service on staging finishes with no connection-wait timeouts and with the total open connections under the database's limit. You receive the connection budget, the change and monitoring queries.

Check whether this job fits

These questions separate the common causes of the same error. They need no access.

Which message do you see?
Do you know how many processes and workers run the service?
Are many connections idle in a transaction at the busy time?

Answer the questions to see whether this job fits.

Nothing is sent anywhere until you choose to email us.

Send an enquiry about this outcome

Checks you can run yourself

  1. Count connections by state (PostgreSQL)

    Ask your engineer to run this read-only query at a busy time. It changes nothing.

    SELECT state, count(*) FROM pg_stat_activity GROUP BY state ORDER BY count(*) DESC;

    Look for: A large count of idle or idle in transaction connections next to the server's limit. Send the counts and the limit.

What you get

  • A connection budget table: processes, workers, pool size per process and the total against the server limit
  • The settings and code change as a pull request
  • Load-test results before and after, with connection counts by state
  • Read-only monitoring queries and the thresholds to alert on

Included

  • One service that connects to one PostgreSQL or MySQL database, built on SQLAlchemy, Django, Rails or a JVM pool
  • Count connections by process and state, and find whether the cause is sizing, leaks, process multiplication, pooler mode or connections held across slow calls
  • Correct the pool settings and any code that holds connections too long, and write the connection budget across all processes
  • A load test before and after at the concurrency you name

Not included

  • Changing the database server's connection limit or hardware in production
  • Installing or operating a pooler in production; a pooler can be tested on staging only
  • More than one service, or a rewrite of the data-access layer
  • Slow queries that hold connections only because they are slow
  • A guarantee for traffic beyond the load-test level

How we know it’s done

Agreed with you before work starts. Each check produces evidence you keep.

  1. In the load test at the agreed concurrency on staging, the service completes with zero connection-wait timeouts and zero server too-many-connections errors.

    Evidence: The load-test report with error counts and the connection counts by state over the run.

  2. The connection budget shows the total of every process type times its pool size, and that total sits below the database limit with the headroom you agree.

    Evidence: The budget table, confirmed in writing by your engineer.

  3. The same load test run on the original configuration reproduces the failure, so the pass is meaningful.

    Evidence: The baseline load-test report.

  4. The monitoring queries return connection counts by state on staging and you have the thresholds to alert on.

    Evidence: Query output from the final run and the written thresholds.

Sign-off. You read the budget and both load-test reports, repeat the test if you wish, and sign off in writing. Payment follows sign-off.

If it fails. If the load test still shows connection-wait failures at the agreed concurrency, you do not pay for this fixed scope and keep the budget and findings. If the cause is a larger architecture limit, we say so and stop.

When it fits, and when we stop

It fits when

  • The service uses SQLAlchemy, Django, Rails or a JVM connection pool, with its version known
  • A staging environment can run the service against a staging database and a load generator
  • You can say how many processes and workers run the service in production
  • An engineer on your side reviews and deploys the change

We stop and tell you if

  • The service cannot run on staging against a realistic database
  • The connection pressure comes from several services we are not given, not the one named
  • The only cure is a larger database server or a different architecture
  • Real customer data would be needed to reproduce the load

What could go wrong

The change is a pull request with its settings in version control. Reverting the commit restores the earlier pool settings. We change nothing on your database server or in production.

Scroll the table sideways to read it all.

RiskHow we handle it
Shrinking a pool starves a busy request type.The load test includes every agreed request type and reports queueing time, not just errors.
A pooler in transaction mode breaks session features the code relies on.Any pooler is tested with the service on staging, and unsupported features are listed before a recommendation.
The total across hosts is miscounted.The budget lists every process type and host count, which your engineer confirms in writing.

A second reviewer checks the connection budget arithmetic against every process type, confirms the load test is the same before and after, and reviews any code that now returns a connection earlier for correctness. Your engineer deploys.

How we deliver

We arrange the work and independent review, then show you the result against the agreed checks. You keep authority over your systems.

  • Agree the service, the concurrency target and the load profile in writing
  • Record the baseline: errors and connection counts by state under the load test on staging
  • Work out the connection budget across every process, and locate leaks and connections held across slow calls
  • Correct the pool settings and code; if a pooler is proposed, test the service against it on staging first
  • Repeat the load test and compare errors and connection counts with the baseline
  • Independent review of the budget, diff and results, then hand over the pull request and monitoring queries

This is a one-off job, not emergency cover or a subscription. We confirm eligibility, the total price, a start window and a delivery date before you accept. Work starts only after agreed inputs, secure access and any permissions are in place. Hosting and database charges are excluded unless the written quote includes them. No charge or booking is created by an enquiry.

Need to keep it working?

Discuss a monthly database health review if you want connection counts and long transactions watched.

Ongoing work is separately scoped and quoted: no monitoring, response-time guarantee or automatic subscription is included in this job.

Explore an ongoing engineering lane, or mention the responsibility you need in your enquiry.

What you can check

This is a new service. We have not delivered this job for a client yet.

Other ways to get this done

  • SQLAlchemy documents the default QueuePool limits, pool_pre_ping, pool_recycle and the need for a fresh pool in each forked process. docs.sqlalchemy.org
  • The HikariCP wiki argues for small pools and tuning by test rather than by a fixed formula. github.com
  • PgBouncer documents its three pooling modes and the session features that do not work in transaction pooling. www.pgbouncer.org

Questions

Can I just raise max_connections?

That raises memory use on the server and can hide a leak or a multiplied pool. We first find the cause, then budget connections against the limit.

Do you install PgBouncer for us?

Not in production. We can test your service against a pooler on staging and list what breaks; you decide and deploy.

My service is not in Python or Ruby.

A JVM service using a standard pool can fit. Other stacks need their own scope; say which in your enquiry.

Send an enquiry

Send us

  • The framework and version, the process model (processes, threads, workers) and the pool settings you know
  • The exact error text, with any addresses removed, and when it happens
  • The database engine, its connection limit and how many connections are open at the busiest time
  • Do not send credentials, a connection string or real data in the first enquiry

Later, once you agree

  • The service code and configuration through an authorised company-controlled repository route
  • A staging environment and database with synthetic data, and the load profile to apply
  • Read-only monitoring figures for connections by state at busy times, exported by your engineer
  • The concurrency target and who deploys the change

You own the database, the code and every credential. We work on a copy you prepare, such as a staging database restored from a backup with personal data removed or replaced, and we hand work back as a pull request or a script. We never ask for production passwords, and we do not connect to your production database. You apply any change to production yourself, after a restore point that you have tested.

A public HTTPS link only, without login details, query strings or fragments. No code or logs.

Sending emails your enquiry and contact address to our team through our mail provider (Resend). It is not kept in a website database. Do not send passwords, keys, recovery links, confidential code or customer records. Your contact email is unverified; nothing is ordered, charged or reserved. Privacy notice.

Email fallback: open your mail app

If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “backend-connection-pool-exhaustion-repaired” as the subject.