Synthetic Industry

Project worker-background-jobs-made-reliable-for-one-service · revised 11 October 2026

Project

Make one service's background jobs and scheduled tasks reliable, and watched

A defined list of job and schedule problems in one service: each job gets a written, tested repeat policy and each schedule a tested missed-run alert, returned as accepted, tested changes.

This asks for a proposal by email. Nothing is charged, and nothing starts, until you have agreed the list, the price and the terms in writing.

The result you are buying

A service grows a collection of background jobs and scheduled tasks, each written separately. Some repeat their side effect when delivery repeats, some stop without anyone knowing, and none of them reports to a person when it fails. Fixing the loudest one leaves the rest as they were, and there is no single list of what runs, who owns it and who is told when it stops.

Who it’s for: An engineering manager or founder whose service relies on background jobs and scheduled tasks that sometimes repeat, stall or silently stop.

Usually starts when: Customers have seen repeated emails or records, a nightly task has gone missing, or a queue backed up unnoticed, and the team wants all of it fixed once rather than job by job.

The result: You buy the finished list for one service. Every agreed job follows a written, tested policy for repeats (done once where the effect can be checked, a stated policy where it cannot), every agreed schedule has a tested missed-run alert to a named person, and a register records what runs, who owns it and who is told.

How the work fits together

The project is complete when every item on the agreed list has been accepted by you or removed from the list in writing, and the closing check with deliberate failures has raised its alerts. You accept the finished list, not each internal step.

  1. Stop one background job doing its work twice, with an idempotency check and replay test Job Included

    One background job that sometimes repeats its side effect gets a database-enforced idempotency record with a claim state and a lease, and a test that replays it through failures.

    Up to three job types on the agreed list

  2. Find why a scheduled job stopped, fix it, and get an alert if it ever misses a run Job Included

    One cron job that silently stopped is diagnosed and fixed, overlapping runs are guarded, and a missed-run alert is wired up and shown to reach a named person.

    Up to four scheduled tasks on the agreed list

  3. Merge duplicate records in one table and add the constraint that stops new ones Job Optional

    Duplicate rows in one PostgreSQL or MySQL table are merged on a copy, with removed rows archived, a reverse script that is run, a reconciliation report and a non-blocking uniqueness guard.

    One table, if you add the clean-up of earlier duplicates

    After: Make one job idempotent

  4. Keep your scheduled jobs and job queues watched, month after month Standing service Optional

    Named cron jobs and queues get heartbeat checks and queue alarms; a missed run, backlog or dead-lettered message reaches a named person, and up to three alerts a month are investigated.

    To keep the result: a monthly service that watches the jobs and queues after the project ends.

How an engagement works

The price covers the agreed list only. Items added later, or found by the inventory beyond the list, are quoted separately, and items found to be bigger jobs are re-scoped with you.

How it starts

  1. You describe the incidents and send the queue system, framework, a rough count of jobs and schedules, and what each job does to the outside world as far as you know, with no code or data.

  2. From that description alone we propose the list of items (their tests, the jobs and schedules in scope) and a fixed price for that list, and the terms.

  3. You agree the list, price and terms in writing. Nothing starts, and we have no access to your code, queues or accounts, before then.

  4. After you agree, the inventory is the first step: you give us a repository or fork and a staging queue, and we inventory the jobs and schedules and their side effects. Any job, schedule or effect that is not on the agreed list is sent back to you and priced in writing before any work on it.

  5. We work through the list, sending a pull request with its test for each item to accept.

  6. We run the closing check with deliberate failures, send the final summary and hand over anything left open.

Who decides what

You decide what is on the list and accept each item. Your team merges and deploys on its own gates and owns the monitoring accounts.

Handover

Each accepted item arrives as a pull request with its test and evidence. Anything left unfinished is handed back with notes.

Sharing your product safely. Send the incidents and versions only. After you agree the project in writing, invite us to a repository or fork you control and provide a staging queue with synthetic data, never real customer data.

What is included, and what is not

  • The inventory and register: each job and schedule, its owner, its alert route and its repeat policy
  • A pull request for each accepted item, with its replay or alert test
  • Heartbeat checks for the scheduled tasks in accounts you own, with test alerts shown
  • A final summary: items accepted, removed or found to need a bigger job

Included

  • One service, its job queue (Sidekiq, Celery or Amazon SQS) and its cron schedules
  • The inventory as the first agreed step: the jobs and schedules, with the side effect of each and what happens if it runs twice or not at all
  • Each job on the list gets an idempotency record with a claimed or completed state and a lease, and a written policy for the case where it cannot tell whether the effect happened, as in the single-job fix
  • A written list of items, quoted from your description and confirmed by the inventory, each with an acceptance test, and a fixed price for the agreed list; anything the inventory adds is priced in writing before work on it
  • A closing check with deliberate failures, and a final summary of each item and how it ended

Not included

  • Cleaning up duplicates already created, unless you add a reconciliation item to the list
  • Queue backlog and dead-letter alarms; the monthly watching service covers those
  • Changing queue infrastructure or moving to another queue product
  • Systemd timers, container schedulers and cloud schedulers
  • More than one service
  • Round-the-clock paging or a response-time guarantee

How we know it’s done

Agreed with you before work starts. Each check produces evidence you keep.

  1. Every item on the agreed list is either accepted by you or removed from the list in writing.

    Evidence: The final summary, with your acceptance or removal recorded for each item.

  2. Each job on the list follows its written policy when its message is delivered more than once, in the replay tests on staging: where the effect is a row in your database or a call that accepts a reference or can be looked up, it happens once; for any other effect, the count of effects is what the item's policy states in advance.

    Evidence: Replay test output attached to each item, with the policy it was tested against.

  3. A deliberately failed run and a deliberately skipped run of each scheduled task on the list each raise an alert to the named person.

    Evidence: Alert messages with timestamps for each task.

  4. The register lists every job and schedule on the list with its owner, alert route and thresholds.

    Evidence: The register, which you can compare with your own inventory.

Sign-off. You accept each item, or send it back with comments, and then you accept the finished list.

If it fails. An item we cannot bring to its test is named with the reason and what it would take, and is taken off the list with a matching change to the price. Nothing is billed as delivered that you have not accepted.

When it fits, and when we stop

It fits when

  • The service uses Sidekiq, Celery or an Amazon SQS consumer, and cron on a Linux host
  • The jobs can run on staging with a test queue and synthetic data
  • The service has a relational database where an idempotency record with a unique constraint can be added
  • If you add the clean-up of earlier duplicates: PostgreSQL or MySQL with InnoDB, a tested restore point, a restored copy with personal data replaced, and one read-only count of duplicate groups that your engineer runs on production and on the copy and sends us
  • You create the heartbeat checks and the alert route in accounts you own, and name who is told
  • A named person can accept each item and merge pull requests

We stop and tell you if

  • Most side effects go to outside systems that offer no reference and no lookup, and you cannot choose a policy for the uncertain case
  • The jobs cannot run on staging without live external calls
  • No one can be named to receive alerts
  • The list is really a rebuild of the job system, so we propose a different scope

What could go wrong

Every item is a separate pull request, so any one can be reverted on its own. Monitors and thresholds live in accounts you own and can be paused or deleted. Nothing is applied to production by us.

Scroll the table sideways to read it all.

RiskHow we handle it
A job is made safe to repeat but its outside effect cannot be detected.The inventory, the first agreed step, marks such effects before any work on the job, and the item states the written policy for the uncertain case (hold for a person, a rare duplicate or a rare miss) instead of promising once-only behaviour.
Alerts are noisy and get muted.Grace times and thresholds are agreed with you and checked in the closing test.
The list grows while we work.The price covers the agreed list. Additions are written in and quoted, never absorbed silently.

Each change is reviewed separately from the work that produced it, with attention to the order of claim, effect and record in every job. You accept every item. No human supervisor is included unless your proposal names one. At launch the work is largely automated, and we say so.

Stays with a person

  • You accept each item
  • You merge, deploy and apply scripts on your own gates, after a tested restore point

Access we would need

  • Read access to a repository or fork you control; no production access
  • Staging copies with personal data removed or replaced

Questions

Do I buy tickets or a result?

A result for one service: an agreed list of jobs and schedules, each tested, and alerts that were seen to fire in a closing check.

What does the project add to buying the jobs one by one?

The single jobs fix one job or one schedule each. The project adds one inventory and register of everything the service runs, one closing check with deliberate failures across the whole list, and one acceptance of the finished list. The starting price is the included jobs at their own starting prices plus that inventory, closing check and summary, so the saving is in coordination, not price. The price is a hypothesis we want to test.

What if a job calls a system that cannot detect repeats?

We say so in the inventory, the first agreed step, and before any work on that job. The item then states a written, tested policy for the uncertain case, and you choose between holding for a person, a rare duplicate and a rare miss, instead of us promising once-only behaviour.

Can the project include the clean-up of old duplicates?

Yes, as an optional item on one table, once the jobs stop creating new duplicates. It is priced into the list.

Send an enquiry

Send us

  • The queue system, the language and framework, a rough count of jobs and scheduled tasks, and what each job does to the outside world as far as you know
  • The incidents so far: repeated effects, missing runs and backlogs, in plain words
  • Who owns the service and who should receive alerts
  • Do not send credentials, code or customer data in the first enquiry

Later, once you agree

  • Only after you have agreed the list, price and terms in writing: a repository or fork you control, with run instructions and the tests
  • A staging queue and database with synthetic data
  • The heartbeat checks you create in accounts you own, with a user limited to that project invited for us or the settings applied by you from our written steps
  • A named person to accept each item and the recipients for alerts

Your databases, queues, code and credentials stay yours. Before you agree in writing we need no code, queue or access. After that we work on staging copies you prepare, with personal data removed or replaced, and hand every change back as a pull request or script. We never ask for production passwords, and we do not connect to production. You apply changes after a restore point you have tested.

A public HTTPS link only, without login details, query strings or fragments. No code or logs.

Sending emails your enquiry and contact address to our team through our mail provider (Resend). It is not kept in a website database. Do not send passwords, keys, recovery links, confidential code or customer records. Your contact email is unverified; nothing is ordered, charged or reserved. Privacy notice.

Email fallback: open your mail app

If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “worker-background-jobs-made-reliable-for-one-service” as the subject.