Testing a single API endpoint is easy. Testing a workflow that spans five services — where step one's output feeds step two, which triggers step three, which writes to a queue that step four consumes — is where most test suites quietly give up.
The usual answer is a big pile of hand-written scripts: one that sets up state, one that walks the sequence, a dozen more that assert on each hop, plus mock servers you maintain by hand for every dependency. It works right up until the workflow changes, at which point you're editing fifteen files to reflect one product change.
I wanted distributed workflow coverage without that maintenance pile. Here's what worked.
Why hand-written scripts don't scale for workflows
A distributed workflow test has two hard parts:
- The sequence — create → update → trigger → consume, in order, with state carried between steps.
- The dependencies — every service in the chain calls databases, queues, and other services that have to behave predictably for the test to be deterministic.
Hand-writing that means scripting each step and standing up (or manually mocking) every dependency. Both grow with the system, and both break when the system changes.
Record the whole sequence instead of scripting it
The approach that removed the pile: instead of scripting the workflow, Keploy records the real request/response sequences as they actually happen across services, and captures the dependencies alongside them.
Because it records the actual flow, multi-step lifecycles — create, mutate, delete — are preserved as connected cases, not reassembled from isolated endpoint checks. And because it auto-generates deterministic mocks for the downstream calls, the workflow replays the same way every time without you standing up the full system or hand-writing a single mock server.
Under the hood it uses eBPF to capture traffic at the network layer, so this works regardless of language — the services in our chain weren't all the same stack, and it didn't care.
What replay looks like
When the workflow changes, you replay the recorded sequence against the new code. It runs the steps in order, against the captured (mocked) dependencies, and compares each response to the baseline. If step three now returns something step four can't handle, the replay flags it — without you maintaining the script that used to test that seam.
keploy test -c "docker compose up"
Run it in CI and distributed-workflow regressions get caught on the pull request.
Where the honesty lives
This isn't a silver bullet for every kind of distributed test. Load and performance characteristics of a workflow are a different problem — reach for k6 or JMeter there. And you still review the recorded flows and decide which ones matter; recording reduces the scripting, it doesn't remove your judgment about coverage.
But for the specific pain of "we have complex multi-service workflows and I refuse to maintain a hundred brittle scripts to test them," recording the real sequences and letting the dependencies be mocked automatically has been the most maintainable approach I've used.
It's open source if you want to try it against your own workflows, and this comparison of API testing tools sets it next to the alternatives by use case. I'd genuinely like to hear how other teams are testing distributed flows — the script-pile approach can't be the best we've got.