One of our workflows runs nightly and processes ~5000 patient records. As of about a week ago it started failing partway through with a timeout. We didn't change the workflow. What's going on?
Hi Kenji — the 30-minute max workflow runtime is a hard platform limit, and your workflow is currently running for ~30m 12s before being killed. So it's not a transient failure — it's a constraint your workflow is now bumping into.
What probably changed: you mentioned 5000 records. If that grew (say from 4000 to 5500), your runtime stretched past 30 minutes. Looking at the run history, your average run time has been climbing for ~2 weeks.
Three options:
Option 2 is usually the lowest-effort if the per-record work is independent. Want me to help model what that would look like?
Per-record work is fully independent (each record is a separate API call). Yes please walk me through Parallel Loop.
Great, that's the easy case. In your workflow editor, replace the current Loop step with a Parallel Loop step. Configure concurrency to 10 (the max). Helix will run up to 10 iterations simultaneously, so 5000 records at 0.4s each goes from ~33 minutes sequential to ~3.3 minutes parallel.
One caveat: your downstream API may rate-limit you. If the API you're hitting has a per-second cap, parallelism can blow through it. Check your downstream's limits before bumping concurrency to the max.
Done. Test run finished in 3m 47s. Going to monitor for a few nights. Thanks Marcus.
Customer's nightly workflow grew past the 30-minute platform max-runtime limit as record volume increased. Switched from sequential Loop to Parallel Loop with concurrency=10. Runtime dropped from ~33min to ~4min. Resolved.