Concurrency and parallelism in LlamaIndex
· 11 min read
Most write-ups about concurrency in LlamaIndex tell you half the story. They show you how to decorate a step with num_workers=5 and call it a day. What they don't tell you is that the storage layer under those parallel steps fails in opposite ways. The default 4 workers guarantee you'll hit both failure modes eventually. I’ve timed the workflow side on llama-index-workflows 2.22.2 (PyPI release) and I’ve lived through the silent corpus inflation and lock-contention meltdowns that happen when those workers hammer the vector store. Here’s the full picture.
