Skip to content

deep_research and financial_report templates instantiate workflow = create_workflow() at module scope — shared state leaks across concurrent requests #724

Description

@Fr3ya

Bug

The official LlamaIndex starter templates deep_research and financial_report both end with a module-level singleton:

# last line of the file
workflow = create_workflow()

create_workflow() builds a Workflow subclass whose __init__ stores per-call state as instance attributes:

# packages/create-llama/templates/components/use-cases/python/deep_research/workflow.py
class DeepResearchWorkflow(Workflow):
    def __init__(self, ...):
        self.context_nodes = []
        self.memory = SimpleComposableMemory.from_defaults(...)
        # ... self.user_request set later in prepare step

LlamaIndex's Workflow.run(...) permits multiple concurrent calls on the same instance, but every step in these templates writes through self.xxx. Two concurrent requests share self.memory, self.context_nodes, and self.user_request — tenant A's query, retrieved doc nodes, and chat history are visible to tenant B's response, and vice versa.

Any project scaffolded via npx create-llama … that picks these templates and deploys with even a single uvicorn worker handling async concurrency (the default in llamaindex-server) inherits the leak.

Affected files (HEAD 97a7d9bc)

File Module-level singleton Instance attrs touched per step
packages/create-llama/templates/components/use-cases/python/deep_research/workflow.py line 588 workflow = create_workflow() self.memory (138-140), self.context_nodes (138, 180), self.user_request (149)
packages/create-llama/templates/components/use-cases/python/financial_report/workflow.py line 330 workflow = create_workflow() self.memory, self.user_request

The sibling files under packages/create-llama/templates/components/agents/python/…/workflows/ define the same Workflow subclass with the same instance-attr layout (the use-cases/ copies just bake in the module-level singleton on top).

Reproducer

We replicate the exact pattern (Workflow subclass + __init__ instance attrs + module-level singleton) using llama-index-core==0.14.22 — the same library version the templates target. The full template can't be imported standalone (heavy LlamaCloud / pgvector / spacy deps), but the bug is in the pattern, not in any specific step body, so a faithful replication is sufficient:

# repro_real.py
import asyncio
from llama_index.core.workflow import (
    Workflow, Context, StartEvent, StopEvent, Event, step,
)

class StartReq(StartEvent):
    user_msg: str

class RetrievedEvent(Event):
    pass

# Mirrors DeepResearchWorkflow's instance-attr layout (workflow.py:138-149)
class DeepResearchWorkflow(Workflow):
    def __init__(self, **kw):
        super().__init__(**kw)
        self.context_nodes = []
        self.memory = []
        self.user_request = None

    @step
    async def prepare(self, ctx, ev: StartReq) -> RetrievedEvent:
        self.user_request = ev.user_msg                              # line 149
        self.memory.append({"role": "user", "content": ev.user_msg}) # lines 158-165
        self.context_nodes.extend([f"doc_for_{ev.user_msg}"])        # lines 179-180
        await asyncio.sleep(0.05)
        return RetrievedEvent()

    @step
    async def finalize(self, ctx, ev: RetrievedEvent) -> StopEvent:
        return StopEvent(result={
            "saw_user_request":  self.user_request,
            "saw_memory":        list(self.memory),
            "saw_context_nodes": list(self.context_nodes),
        })


workflow = DeepResearchWorkflow(timeout=10)   # mirrors the template's line 588

async def main():
    a, b = await asyncio.gather(
        workflow.run(user_msg="TENANT_A_secret_aaaa"),
        workflow.run(user_msg="TENANT_B_unrelated_bbbb"),
    )
    print("Run A saw:", a)
    print("Run B saw:", b)

asyncio.run(main())

Output:

Run A saw: {'saw_user_request': 'TENANT_B_unrelated_bbbb',
            'saw_memory':        [{'role':'user','content':'TENANT_A_secret_aaaa'},
                                  {'role':'user','content':'TENANT_B_unrelated_bbbb'}],
            'saw_context_nodes': ['doc_for_TENANT_A_secret_aaaa',
                                  'doc_for_TENANT_B_unrelated_bbbb']}
Run B saw: {'saw_user_request': 'TENANT_B_unrelated_bbbb',  ...same merged state...}

Both concurrent .run() calls observe each other's user_request, memory, and context_nodes. The "last-writer wins" race on self.user_request plus the share-and-append on self.memory / self.context_nodes is a cross-tenant data leak.

Expected behavior

Either:

  • The templates instantiate a fresh Workflow per request (move the singleton out, document create_workflow() as the per-request factory). Llama-index-server's WorkflowFactory already has the right API for this, but the template ships the singleton.
  • Or the templates use Workflow's Context.set/get (which is per-run) instead of self.xxx instance attrs for memory, context_nodes, and user_request.

Both fixes preserve correctness; the first is the smaller diff (one line per template).

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions