Skip to main content

A2A's task lifecycle: what the async contract fixes, and what it leaves to you

· 10 min read
Vadim Nicolai
Senior Software Engineer

The A2A v1.0 specification says SendMessage "MUST return immediately with either task information or response message" (§3.1.1). A dozen subsections later it says "Operations are blocking by default", and a blocking call "MUST wait until the task reaches a terminal state ... or an interrupted state" (§3.2.2). The default client does not wait for completion. It waits for a pause, which may be a human deciding whether to approve something.

That contradiction is a fair summary of the whole protocol: a precise vocabulary for where work is, and almost no guarantee that it happened.

Nine states, one split that matters​

TaskState has nine values. The proto labels four "terminal" and two "interrupted"; the rest carry no label.

ClassStatesMeaning for the client
TerminalTASK_STATE_COMPLETED, TASK_STATE_FAILED, TASK_STATE_CANCELED, TASK_STATE_REJECTEDEnded. Any continuation is a new task.
InterruptedTASK_STATE_INPUT_REQUIRED, TASK_STATE_AUTH_REQUIREDPaused. Continue the same task.
Active / unlabeledTASK_STATE_UNSPECIFIED, TASK_STATE_SUBMITTED, TASK_STATE_WORKINGIn flight or unknown.

Terminal is enforced. A message to a terminal task returns UnsupportedOperationError (§3.1.1), mapped to JSON-RPC -32004, gRPC FAILED_PRECONDITION, HTTP 400 (§5.4). The topic guide states it plainly: "Once a task reaches a terminal state (completed, canceled, rejected, or failed), it cannot restart", and refinements "must initiate a new task within the same contextId" (Life of a Task).

Interrupted is resumable: "the client continues the interaction by sending a new message with the same taskId and contextId" (§3.4.3). Mismatched taskId/contextId pairs MUST be rejected.

Rejection is not admission control. REJECTED "may be done during initial task creation or later once an agent has determined it can't or won't proceed" (proto TaskState). A task that was returned as WORKING can still be rejected later, so a successful send is not acceptance.

Retry belongs entirely to the caller​

Three rules combine:

  • Task ids are server-generated; "Client-provided taskId values for creating new tasks is NOT supported" (§3.4.2).
  • Terminal tasks never restart.
  • TaskStatus has exactly three fields: state, message, timestamp (§4.1.2). No retryable flag, attempt count or backoff hint.

So a retry is a new task with a new id, and nothing on the wire links it to the attempt it replaces unless you put the old id in referenceTaskIds, which clients SHOULD use for related tasks (§3.4.3). The partner agent does not know a retry is happening.

Duplicate suppression is optional: SendMessage "MAY be idempotent. Agents may utilize the messageId to detect duplicate messages" (§3.3.1). The Agent Card capabilities do not advertise whether an agent dedups.

Now combine that with the blocking default. A blocking call held open through INPUT_REQUIRED meets an idle-timeout proxy, the client sees what looks like a request that never landed, and resends. Server-assigned ids plus MAY-level dedup means two tasks, two sets of side effects. If the work was a payment, the duplicate is a refund.

Cancellation is honest and incomplete. "Success is not guaranteed", a terminal task returns TaskNotCancelableError (-32002) (§3.1.5, §5.4), and cancel is idempotent except that a repeat "MAY return TaskNotFoundError if the task has already been canceled and purged" (§3.3.1). There is no compensation hook. CANCELED and FAILED describe the record, not the world: a task that reserved inventory and then failed has no protocol-level place to report the release. That is the problem sagas (Garcia-Molina and Salem, 1987) were invented for, and A2A leaves it to you.

Three channels, three places to park the burden​

ChannelOperationsSpec's own trade-off (§3.5.1)Capability gate
PollingGetTask"Higher latency, potential for unnecessary requests"; best behind restrictive firewallsnone
StreamingSendStreamingMessage, SubscribeToTask"Requires persistent connection support"capabilities.streaming, else UnsupportedOperationError
PushCreateTaskPushNotificationConfig + webhook"client must be reachable via HTTP"capabilities.pushNotifications, else PushNotificationNotSupportedError (-32003)

Gates and error codes: §3.3.4, §5.4. Method names are the v1.0 JSON-RPC/gRPC names (§5.3); the REST forms are POST /message:send, GET /tasks/{id}, POST /tasks/{id}:subscribe.

Streaming has the strongest semantics. Events "MUST NOT be reordered", every concurrent stream receives "the same events in the same order", and "The task lifecycle is independent of any individual stream's lifecycle" (§3.5.2). SubscribeToTask MUST send the current Task first, which "prevents a potential loss of information between a call to GetTask and calling SubscribeToTask", and fails on terminal tasks (§3.1.6). Artifact chunks carry append and lastChunk (§4.2.2).

Guard against one discrepancy. §3.1.2 says the stream MUST close on the four terminal states. The streaming topic page says it closes on "a terminal or interrupted state (e.g., ... INPUT_REQUIRED)". Your reader must handle both: a stream that ends at INPUT_REQUIRED and one that stays open through it. The spec names a2a.proto "the single authoritative normative definition" (§1.4); treat numbered sections as norm and topic pages as intent.

Push is an attempt, not a delivery​

The obligation is one word wide: "Agents MUST attempt delivery at least once for each configured webhook" (§4.3.3). The same section says agents MAY retry with exponential backoff; §13.2 says SHOULD. Both allow stopping "after a configured number of consecutive failures", and both recommend a 10-30 second webhook timeout. Implement the SHOULD: it satisfies the MAY, never the reverse.

No dead-letter queue, no replay operation. The config "MUST persist until task completion or explicit deletion" (§3.1.7), nothing longer. A receiver that was down when a task completed has no protocol way to learn the event existed. It has to ask.

The receiver side is not free either (§13.2): payloads are a StreamResponse with exactly one of task, message, statusUpdate, artifactUpdate, sent as application/a2a+json over plain HTTP whatever binding the agent speaks; clients MUST answer 2xx and SHOULD process idempotently "as duplicate deliveries may occur". Agents SHOULD reject webhook targets in 127.0.0.0/8, 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, localhost and link-local, because an open webhook makes the agent an SSRF vector or a "DDoS amplifier" (streaming topic).

Events are hints, GetTask is truth​

The topic guide describes the pattern: on a push, the client "typically uses the GetTask RPC method ... to retrieve the complete, updated Task object" (streaming topic). §3.7 backs it normatively: results SHOULD be Artifacts, not Messages; a reconnecting stream "MAY not receive all status update messages"; and "Messages MUST NOT be considered a reliable delivery mechanism for critical information".

The reconciliation sweep is ListTasks (§3.1.4, proto). Know its defaults:

FieldBehaviour
pageSizedefault 50, min 1, max 100
pageToken / nextPageTokencursor-based; final page returns ""
orderingstatus timestamp, descending
statusTimestampAfterinclusive (>=), so overlapping sweeps re-see boundary tasks
includeArtifactsdefault false; artifacts are omitted entirely
scopingMUST return only tasks visible to the caller, even with no filter (§13.1)

A sweep that forgets includeArtifacts returns task shells and looks healthy until someone needs the output. And §13.1 is explicit that "Authorization boundaries are defined by each agent's authorization model, not prescribed by the protocol".

AUTH_REQUIRED is a state, not a permission​

§7.6 gives the canonical example: "human approval before a destructive action is taken". Credentials arrive out of band by default, and the agent "MAY immediately continue Task processing after receiving the credential, without a requirement that clients send a follow-up message" (§7.6.1). A client without an open stream "risks missing Task updates", and a client that is itself an agent can delegate upward, "forming a chain of Tasks" in AUTH_REQUIRED (§7.6.2). An approval queue has to render chains, not rows.

§7.6.4 declines to define "the scope, representation, validity, or revocation semantics" of the grant, and says agents "MUST NOT treat the TASK_STATE_AUTH_REQUIRED state transition, by itself, as authorization". Log the state and the grant separately. A transition from AUTH_REQUIRED to WORKING is not evidence of a permission.

What the spec fixes, what it leaves open​

QuestionSpec's answer
Can a finished task change?No. Terminal is immutable (§3.1.1).
Who retries?Unspecified. No retry fields in TaskStatus.
Are duplicates suppressed?MAY, via messageId (§3.3.1).
What if the webhook is down?One attempt guaranteed; retry MAY/SHOULD; no replay.
Task deadline or TTL?No such field on Task (§4.1.1).
How long do tasks live?Unspecified. TaskNotFoundError covers "invalid, expired, or already completed and purged" (§3.3.2); contexts MAY expire and SHOULD be documented (§3.4.1).
What does an auth grant cover?Implementation-defined (§7.6.4).

Six decisions to write down before the first integration​

  1. Channel per caller class: polling behind restrictive egress, streaming for interactive surfaces, push for server-to-server (§3.5.1).
  2. Set returnImmediately: true on anything that can pause. Never let a request thread wait for a human (§3.2.2).
  3. Publish what messageId does on your agent and what a caller should assume after a timeout (§3.3.1).
  4. Run a ListTasks sweep on statusTimestampAfter, with includeArtifacts when outputs matter, deduping the inclusive boundary (§3.1.4).
  5. Decide what CANCELED and FAILED leave behind and where you record compensation (§3.1.5).
  6. Publish retention, and treat TaskNotFoundError as ambiguous by design; default history length is "implementation-defined" (§3.2.4).

The spec even names the alarms: agents SHOULD monitor "rapid task creation, excessive cancellations" and SHOULD "implement appropriate data retention policies" (§13.4). Rapid task creation is the signature of duplicate submission. Excessive cancellation is the signature of a retry policy nobody owns.

Nine states is an afternoon of reading. The six decisions are the quarter.