A2A

Bring me my Llama

Posted by Isaac on Tuesday, October 6, 2026

Ever since I heard about the Agent to Agent (A2A) protocol, I’ve been curious what it was and how it might be useful.

The general idea of A2A is that it’s a standard that Agentic workloads can leverage to discover and talk to each other using “cards” that advertise capabilities.

We will look at two A2A examples, how they work and test them. Then we’ll try and recreate them with python containerize servers with REST endpoints.

Let’s start with an A2A hello world example.

A Quick Hello World

Let’s start by getting the Samples locally.

$ git clone https://github.com/a2aproject/a2a-samples.git
$ cd a2a-samples

From here we can go to samples/python/agents/helloworld and setup the Python environment

isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/helloworld$ python3 -m venv .venv
isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/helloworld$ source .venv/bin/activate
(.venv) isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/helloworld$ pip install -r requirements.txt
Collecting a2a-sdk==1.1.0 (from -r requirements.txt (line 1))
  Downloading a2a_sdk-1.1.0-py3-none-any.whl.metadata (9.7 kB)
Collecting uvicorn (from -r requirements.txt (line 2))
  Downloading uvicorn-0.54.0-py3-none-any.whl.metadata (6.6 kB)
Collecting pytest (from -r requirements.txt (line 3))
  Using cached pytest-9.1.1-py3-none-any.whl.metadata (7.6 kB)
Collecting sse-starlette (from -r requirements.txt (line 4))
  Downloading sse_starlette-3.5.0-py3-none-any.whl.metadata (15 kB)
Collecting google-api-core>=1.26.0 (from a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Downloading google_api_core-2.40.0-py3-none-any.whl.metadata (3.2 kB)
Collecting googleapis-common-protos>=1.70.0 (from a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Downloading googleapis_common_protos-1.75.5-py3-none-any.whl.metadata (8.5 kB)
Collecting httpx-sse>=0.4.0 (from a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Downloading httpx_sse-0.4.3-py3-none-any.whl.metadata (9.7 kB)
Collecting httpx>=0.28.1 (from a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached httpx-0.28.1-py3-none-any.whl.metadata (7.1 kB)
Collecting json-rpc>=1.15.0 (from a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Downloading json_rpc-1.15.0-py2.py3-none-any.whl.metadata (7.2 kB)
Collecting packaging>=24.0 (from a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached packaging-26.3-py3-none-any.whl.metadata (3.5 kB)
Collecting protobuf<7,>=5.29.5 (from a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Downloading protobuf-6.33.6-cp39-abi3-manylinux2014_x86_64.whl.metadata (593 bytes)
Collecting pydantic>=2.11.3 (from a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached pydantic-2.13.5-py3-none-any.whl.metadata (110 kB)
Collecting click>=7.0 (from uvicorn->-r requirements.txt (line 2))
  Using cached click-8.5.0-py3-none-any.whl.metadata (2.6 kB)
Collecting h11>=0.8 (from uvicorn->-r requirements.txt (line 2))
  Using cached h11-0.16.0-py3-none-any.whl.metadata (8.3 kB)
Collecting iniconfig>=1.0.1 (from pytest->-r requirements.txt (line 3))
  Using cached iniconfig-2.3.0-py3-none-any.whl.metadata (2.5 kB)
Collecting pluggy<2,>=1.5 (from pytest->-r requirements.txt (line 3))
  Using cached pluggy-1.6.0-py3-none-any.whl.metadata (4.8 kB)
Collecting pygments>=2.7.2 (from pytest->-r requirements.txt (line 3))
  Using cached pygments-2.21.0-py3-none-any.whl.metadata (2.5 kB)
Collecting starlette>=0.49.1 (from sse-starlette->-r requirements.txt (line 4))
  Downloading starlette-1.7.0-py3-none-any.whl.metadata (6.6 kB)
Collecting anyio>=4.7.0 (from sse-starlette->-r requirements.txt (line 4))
  Using cached anyio-4.15.1-py3-none-any.whl.metadata (4.7 kB)
Collecting idna>=2.8 (from anyio>=4.7.0->sse-starlette->-r requirements.txt (line 4))
  Downloading idna-3.20-py3-none-any.whl.metadata (7.2 kB)
Collecting typing_extensions>=4.16.0 (from anyio>=4.7.0->sse-starlette->-r requirements.txt (line 4))
  Using cached typing_extensions-4.16.0-py3-none-any.whl.metadata (3.3 kB)
Collecting proto-plus<2.0.0,>=1.26.1 (from google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Downloading proto_plus-1.29.0-py3-none-any.whl.metadata (2.2 kB)
Collecting google-auth<3.0.0,>=2.14.1 (from google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Downloading google_auth-2.59.0-py3-none-any.whl.metadata (6.0 kB)
Collecting requests<3.0.0,>=2.33.0 (from google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached requests-2.34.2-py3-none-any.whl.metadata (4.8 kB)
Collecting opentelemetry-api<2.0.0,>=1.44.0 (from google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Downloading opentelemetry_api-1.45.0-py3-none-any.whl.metadata (1.4 kB)
Collecting pyasn1-modules>=0.2.1 (from google-auth<3.0.0,>=2.14.1->google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached pyasn1_modules-0.4.2-py3-none-any.whl.metadata (3.5 kB)
Collecting cryptography>=41.0.5 (from google-auth<3.0.0,>=2.14.1->google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached cryptography-50.0.1-cp311-abi3-manylinux_2_34_x86_64.whl.metadata (4.3 kB)
Collecting charset_normalizer<4,>=2 (from requests<3.0.0,>=2.33.0->google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached charset_normalizer-3.5.1-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl.metadata (45 kB)
Collecting urllib3<3,>=1.26 (from requests<3.0.0,>=2.33.0->google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Downloading urllib3-2.8.0-py3-none-any.whl.metadata (7.4 kB)
Collecting certifi>=2023.5.7 (from requests<3.0.0,>=2.33.0->google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached certifi-2026.7.22-py3-none-any.whl.metadata (2.5 kB)
Collecting cffi>=2.0.0 (from cryptography>=41.0.5->google-auth<3.0.0,>=2.14.1->google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached cffi-2.1.1-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (2.5 kB)
Collecting pycparser (from cffi>=2.0.0->cryptography>=41.0.5->google-auth<3.0.0,>=2.14.1->google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached pycparser-3.0-py3-none-any.whl.metadata (8.2 kB)
Collecting httpcore==1.* (from httpx>=0.28.1->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached httpcore-1.0.9-py3-none-any.whl.metadata (21 kB)
Collecting pyasn1<0.7.0,>=0.6.1 (from pyasn1-modules>=0.2.1->google-auth<3.0.0,>=2.14.1->google-api-core>=1.26.0->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached pyasn1-0.6.4-py3-none-any.whl.metadata (8.4 kB)
Collecting annotated-types>=0.6.0 (from pydantic>=2.11.3->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached annotated_types-0.8.0-py3-none-any.whl.metadata (15 kB)
Collecting pydantic-core==2.46.5 (from pydantic>=2.11.3->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached pydantic_core-2.46.5-cp314-cp314-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (6.6 kB)
Collecting typing-inspection>=0.4.2 (from pydantic>=2.11.3->a2a-sdk==1.1.0->-r requirements.txt (line 1))
  Using cached typing_inspection-0.4.4-py3-none-any.whl.metadata (2.6 kB)
Downloading a2a_sdk-1.1.0-py3-none-any.whl (241 kB)
Downloading protobuf-6.33.6-cp39-abi3-manylinux2014_x86_64.whl (323 kB)
Downloading uvicorn-0.54.0-py3-none-any.whl (87 kB)
Using cached pytest-9.1.1-py3-none-any.whl (386 kB)
Using cached pluggy-1.6.0-py3-none-any.whl (20 kB)
Downloading sse_starlette-3.5.0-py3-none-any.whl (17 kB)
Using cached anyio-4.15.1-py3-none-any.whl (132 kB)
Using cached click-8.5.0-py3-none-any.whl (125 kB)
Downloading google_api_core-2.40.0-py3-none-any.whl (214 kB)
Downloading google_auth-2.59.0-py3-none-any.whl (262 kB)
Downloading googleapis_common_protos-1.75.5-py3-none-any.whl (307 kB)
Downloading opentelemetry_api-1.45.0-py3-none-any.whl (60 kB)
Downloading proto_plus-1.29.0-py3-none-any.whl (50 kB)
Using cached requests-2.34.2-py3-none-any.whl (73 kB)
Using cached charset_normalizer-3.5.1-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl (251 kB)
Downloading idna-3.20-py3-none-any.whl (69 kB)
Downloading urllib3-2.8.0-py3-none-any.whl (135 kB)
Using cached certifi-2026.7.22-py3-none-any.whl (136 kB)
Using cached cryptography-50.0.1-cp311-abi3-manylinux_2_34_x86_64.whl (4.7 MB)
Using cached cffi-2.1.1-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (221 kB)
Using cached h11-0.16.0-py3-none-any.whl (37 kB)
Using cached httpx-0.28.1-py3-none-any.whl (73 kB)
Using cached httpcore-1.0.9-py3-none-any.whl (78 kB)
Downloading httpx_sse-0.4.3-py3-none-any.whl (9.0 kB)
Using cached iniconfig-2.3.0-py3-none-any.whl (7.5 kB)
Downloading json_rpc-1.15.0-py2.py3-none-any.whl (39 kB)
Using cached packaging-26.3-py3-none-any.whl (129 kB)
Using cached pyasn1_modules-0.4.2-py3-none-any.whl (181 kB)
Using cached pyasn1-0.6.4-py3-none-any.whl (84 kB)
Using cached pydantic-2.13.5-py3-none-any.whl (472 kB)
Using cached pydantic_core-2.46.5-cp314-cp314-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.1 MB)
Using cached annotated_types-0.8.0-py3-none-any.whl (13 kB)
Using cached pygments-2.21.0-py3-none-any.whl (1.3 MB)
Downloading starlette-1.7.0-py3-none-any.whl (78 kB)
Using cached typing_extensions-4.16.0-py3-none-any.whl (45 kB)
Using cached typing_inspection-0.4.4-py3-none-any.whl (14 kB)
Using cached pycparser-3.0-py3-none-any.whl (48 kB)
Installing collected packages: json-rpc, urllib3, typing_extensions, pygments, pycparser, pyasn1, protobuf, pluggy, packaging, iniconfig, idna, httpx-sse, h11, click, charset_normalizer, certifi, annotated-types, uvicorn, typing-inspection, requests, pytest, pydantic-core, pyasn1-modules, proto-plus, opentelemetry-api, httpcore, googleapis-common-protos, cffi, anyio, starlette, pydantic, httpx, cryptography, sse-starlette, google-auth, google-api-core, a2a-sdk
Successfully installed a2a-sdk-1.1.0 annotated-types-0.8.0 anyio-4.15.1 certifi-2026.7.22 cffi-2.1.1 charset_normalizer-3.5.1 click-8.5.0 cryptography-50.0.1 google-api-core-2.40.0 google-auth-2.59.0 googleapis-common-protos-1.75.5 h11-0.16.0 httpcore-1.0.9 httpx-0.28.1 httpx-sse-0.4.3 idna-3.20 iniconfig-2.3.0 json-rpc-1.15.0 opentelemetry-api-1.45.0 packaging-26.3 pluggy-1.6.0 proto-plus-1.29.0 protobuf-6.33.6 pyasn1-0.6.4 pyasn1-modules-0.4.2 pycparser-3.0 pydantic-2.13.5 pydantic-core-2.46.5 pygments-2.21.0 pytest-9.1.1 requests-2.34.2 sse-starlette-3.5.0 starlette-1.7.0 typing-inspection-0.4.4 typing_extensions-4.16.0 urllib3-2.8.0 uvicorn-0.54.0

Now we can fire up the A2A server

(.venv) isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/helloworld$ python __main__.py
INFO:     Started server process [1392384]
INFO:     Waiting for application startup.
INFO:     Application startup complete.
INFO:     Uvicorn running on http://127.0.0.1:9999 (Press CTRL+C to quit)

With the server up, we can go to another window and fire up the client

(.venv) isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/helloworld$ python test_client.py

Starting an internactive session with A2A Server [http://127.0.0.1:9999]
Use `exit` to quit.
user >

It will just repeat back our words as this is a hello world demo app

Let’s take a look at the main.py code

if __name__ == '__main__':
    # --8<-- [start:AgentSkill]
    # Defines the abilities or functions that agent can perform.
    skill = AgentSkill(
        id='echo_bot',
        name='Echo Bot',
        description='An example agent that acknowledges client request and responds with a "Hello World" message.',
        input_modes=['text/plain'],
        output_modes=['text/plain'],
        tags=['a2a', 'echo-example'],
        examples=['hi', 'how are you'],
    )
    # --8<-- [end:AgentSkill]
    # Defines an optional additional skill for the agent that is not visible in the public card.
    extended_skill = AgentSkill(
        id='echo_bot_super_mode',
        name='Echo Bot (Super Mode)',
        description='An extended version of Echo Bot that responds with extra enthusiasm!',
        tags=['a2a', 'echo-example', 'extended'],
        examples=['super hi', 'give me a super hello'],
    )

You can see the skills created. They are like functions that can be exposed later in cards.

Next we see the cards

# --8<-- [start:AgentCard]
# Define a public-facing agent card that allows clients to discover your agent's capabilities.
public_agent_card = AgentCard(
    # Basic identity information of A2A server
    name='Hello World Agent',  # Identity
    description='Just a hello world agent',
    version='0.0.1',
    # Default Media Types for the agent's interactions
    default_input_modes=['text/plain'],  # Supported media types
    default_output_modes=['text/plain'],
    # Supported A2A features (like streaming or extended config)
    capabilities=AgentCapabilities(streaming=True, extended_agent_card=True),
    # Ordered list of endpoints and protocols where the service can be reached
    supported_interfaces=[
        AgentInterface(
            protocol_binding='JSONRPC',
            url='http://127.0.0.1:9999',
            protocol_version='1.0',
        )
    ],
    # The list of AgentSkill objects that this agent offers
    skills=[skill],
    # Optional attributes (omitted here for simplicity):
    # icon_url                         -> A URL to an icon representing the agent
)
# --8<-- [end:AgentCard]

# Defines the authenticated extended agent card with
# extended skills that are visible only to authenticated users
extended_agent_card = AgentCard(
    name='Hello World Agent - Extended Edition',
    description='The full-featured hello world agent for authenticated users.',
    version='0.0.2',
    default_input_modes=['text/plain'],
    default_output_modes=['text/plain'],
    capabilities=AgentCapabilities(streaming=True, extended_agent_card=True),
    supported_interfaces=[
        AgentInterface(
            protocol_binding='JSONRPC',
            url='http://127.0.0.1:9999',
            protocol_version='1.0',
        )
    ],
    skills=[
        skill,
        extended_skill,
    ],  # Both skills for the extended card
)

now this code has no route to the “extended” version, but I’m sure that code could get added to the routes

# --8<-- [start:ServerRoutes]
 # Creating the routes for the A2A server
 # These routes handle the incoming requests from the clients
 # and the outgoing responses to the clients
 routes = []

 # Create routes for the agent card
 routes.extend(create_agent_card_routes(public_agent_card))

 # Create routes for the JSONRPC protocol
 # Alternatively, you can choose GRPC or HTTP_JSON as protocol bindings
 # based on your requirements
 routes.extend(create_jsonrpc_routes(request_handler, '/'))
 # --8<-- [end:ServerRoutes]
 # --8<-- [start:AppServer]

 # Create a web app with the defined routes
 # Here we are using Starlette, a lightweight ASGI web framework to serve the agent
 # Alternatively, you can choose FastAPI or other ASGI frameworks
 app = Starlette(routes=routes)

 # Run the app
 # Uvicorn is a production-ready ASGI HTTP server
 uvicorn.run(app, host='127.0.0.1', port=9999)
 # --8<-- [end:AppServer]

The Client can be a few different things. In this case, we fire up an a2a client and look for an “A2A Card Resolver” on that port (9999)

@pytest.fixture(scope='session', autouse=True)
def start_server():
    server_path = Path(__file__).parent / '__main__.py'
    process = subprocess.Popen(  # noqa: S603
        [sys.executable, str(server_path)],
        stdout=subprocess.PIPE,
        stderr=subprocess.PIPE,
    )
    # Wait a moment for the server to start
    time.sleep(1.5)

    yield

    process.terminate()
    process.wait()


async def get_agent_card():
    print('Initializes the A2ACardResolver instance with an HTTP client')
    # --8<-- [start:A2ACardResolver]
    import httpx  # noqa: PLC0415

    from a2a.client import A2ACardResolver  # noqa: PLC0415

    # Initializes the A2ACardResolver instance with an HTTP client, base URL,
    # and uses the default path for the agent card.
    async with httpx.AsyncClient() as httpx_client:
        resolver = A2ACardResolver(
            httpx_client=httpx_client,
            base_url='http://127.0.0.1:9999',
            # Provide agent_card_path, if your agent uses a different path
            # agent_card_path=''  # noqa: ERA001
        )
        public_agent_card = await resolver.get_agent_card()
        # --8<-- [end:A2ACardResolver]
        print('\nSuccessfully fetched the public agent card:')
    return public_agent_card

It is setup to connect to the standard and extended messages, but only when run with pytest


async def send_message(text_query: str = 'Hi there'):
    public_agent_card = await get_agent_card()
    print('\n--- Public Agent Card - Non-Streaming Call ---')
    # --8<-- [start:message_send]
    from a2a.client import ClientConfig, create_client  # noqa: PLC0415
    from a2a.helpers import new_text_message  # noqa: PLC0415
    from a2a.types import Role, SendMessageRequest  # noqa: PLC0415

    print('\nInitializing a non-streaming client.')
    config = ClientConfig(streaming=False)
    client = await create_client(agent=public_agent_card, client_config=config)

    # Creates a new text message to be sent to the A2A Server.
    # Ex: text_query = 'Why is the sky blue?'  # noqa: ERA001
    message = new_text_message(text_query, role=Role.ROLE_USER)
    request = SendMessageRequest(message=message)

    print('Response:')
    async for chunk in client.send_message(request):
        print(chunk)
    # --8<-- [end:message_send]
    await client.close()


async def send_steaming_message(text_query: str = 'Hi there'):
    public_agent_card = await get_agent_card()

    # Creates a new text message to be sent to the A2A Server.
    # Ex: text_query = 'Why is the sky blue?'  # noqa: ERA001
    message = new_text_message(text_query, role=Role.ROLE_USER)
    request = SendMessageRequest(message=message)

    print('\n--- Public Agent Card - Streaming Call ---')
    # --8<-- [start:message_stream]
    print('\nInitializing a streaming client.')
    client_config = ClientConfig(streaming=True)  # Streaming
    client = await create_client(agent=public_agent_card, client_config=client_config)

    print('Response:')
    async for chunk in client.send_message(request):
        print(chunk)
    # --8<-- [end:message_stream]
    await client.close()


async def show_extended_card():
    from a2a.helpers import display_agent_card  # noqa: PLC0415
    from a2a.types import GetExtendedAgentCardRequest  # noqa: PLC0415

    public_agent_card = await get_agent_card()
    config = ClientConfig(streaming=False)
    client = await create_client(agent=public_agent_card, client_config=config)

    print('\n--- Extended Agent Card - Non-Streaming Call ---')
    extended_card = await client.get_extended_agent_card(GetExtendedAgentCardRequest())
    print('\nSuccessfully fetched the authenticated extended agent card:')
    display_agent_card(extended_card)
    await client.close()


def test_client_workflow(text_query: str = 'Hi there!'):
    """
    This test function is intended to be used to test the client workflow
    for multiple queries
    """
    asyncio.run(show_agent_card())
    asyncio.run(send_message(text_query))
    asyncio.run(send_steaming_message(text_query))
    asyncio.run(show_extended_card())


def main():
    print('\nStarting an internactive session with A2A Server [http://127.0.0.1:9999]')
    print('Use `exit` to quit.')
    prompt = input('user > ')
    while prompt and prompt != 'exit':
        asyncio.run(send_message(prompt))
        prompt = input('--\nuser > ')

Let’s try those out by adding a streaming message instead of async and showing the extended card in our main()

(.venv) isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/helloworld$ git diff ./test_client.py
diff --git a/samples/python/agents/helloworld/test_client.py b/samples/python/agents/helloworld/test_client.py
index b47de00..c2e487f 100644
--- a/samples/python/agents/helloworld/test_client.py
+++ b/samples/python/agents/helloworld/test_client.py
@@ -135,6 +135,8 @@ def main():
     prompt = input('user > ')
     while prompt and prompt != 'exit':
         asyncio.run(send_message(prompt))
+        asyncio.run(send_steaming_message(prompt))
+        asyncio.run(show_extended_card())
         prompt = input('--\nuser > ')

Here you can see both the send_message and send_streaming_message called

(.venv) isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/helloworld$ python test_client.py

Starting an internactive session with A2A Server [http://127.0.0.1:9999]
Use `exit` to quit.
user > Testing more
Initializes the A2ACardResolver instance with an HTTP client

Successfully fetched the public agent card:

--- Public Agent Card - Non-Streaming Call ---

Initializing a non-streaming client.
Response:
task {
  id: "6ccda08c-0e64-433d-9552-c5bc74ff968f"
  context_id: "68e76869-9287-4bfc-8eca-2897dbdc9529"
  status {
    state: TASK_STATE_COMPLETED
    message {
      message_id: "fc5472e9-ff62-4693-bf01-cee22ef6d3af"
      role: ROLE_AGENT
      parts {
        text: "Request is completed!"
      }
    }
    timestamp {
      seconds: 1790939361
      nanos: 995911000
    }
  }
  artifacts {
    artifact_id: "ada9c998-76ae-49e6-b7ba-c775990b0b93"
    parts {
      text: "Hello, World! I have received your request (Testing more)"
      media_type: "text/plain"
    }
  }
  history {
    message_id: "c523b8b6-00d8-4a77-9411-48b674bc43dd"
    context_id: "68e76869-9287-4bfc-8eca-2897dbdc9529"
    task_id: "6ccda08c-0e64-433d-9552-c5bc74ff968f"
    role: ROLE_USER
    parts {
      text: "Testing more"
    }
  }
  history {
    message_id: "a5f4327d-b76f-4871-8cb4-314fcfdda8f8"
    role: ROLE_AGENT
    parts {
      text: "Processing request..."
    }
  }
}

Initializes the A2ACardResolver instance with an HTTP client

Successfully fetched the public agent card:

--- Public Agent Card - Streaming Call ---

Initializing a streaming client.
Response:
task {
  id: "cf694ea0-089b-4118-b3e9-47efa6a4803c"
  context_id: "338cbed6-b120-4608-bdfb-3d8a0930476a"
  status {
    state: TASK_STATE_SUBMITTED
  }
  history {
    message_id: "27be29d8-51aa-44a2-8af9-67f271bc7f90"
    context_id: "338cbed6-b120-4608-bdfb-3d8a0930476a"
    task_id: "cf694ea0-089b-4118-b3e9-47efa6a4803c"
    role: ROLE_USER
    parts {
      text: "Testing more"
    }
  }
}

status_update {
  task_id: "cf694ea0-089b-4118-b3e9-47efa6a4803c"
  context_id: "338cbed6-b120-4608-bdfb-3d8a0930476a"
  status {
    state: TASK_STATE_WORKING
    message {
      message_id: "866f4fba-a28b-49c0-9c02-8898a6960711"
      role: ROLE_AGENT
      parts {
        text: "Processing request..."
      }
    }
    timestamp {
      seconds: 1790939362
      nanos: 19259000
    }
  }
}

artifact_update {
  task_id: "cf694ea0-089b-4118-b3e9-47efa6a4803c"
  context_id: "338cbed6-b120-4608-bdfb-3d8a0930476a"
  artifact {
    artifact_id: "4ace74f8-a912-4faa-9756-5ef51709defe"
    parts {
      text: "Hello, World! I have received your request (Testing more)"
      media_type: "text/plain"
    }
  }
}

status_update {
  task_id: "cf694ea0-089b-4118-b3e9-47efa6a4803c"
  context_id: "338cbed6-b120-4608-bdfb-3d8a0930476a"
  status {
    state: TASK_STATE_COMPLETED
    message {
      message_id: "f40d79cd-48f7-43bb-b8cb-6e127c83f2eb"
      role: ROLE_AGENT
      parts {
        text: "Request is completed!"
      }
    }
    timestamp {
      seconds: 1790939362
      nanos: 19355000
    }
  }
}

Initializes the A2ACardResolver instance with an HTTP client

Successfully fetched the public agent card:

--- Extended Agent Card - Non-Streaming Call ---

Successfully fetched the authenticated extended agent card:
====================================================
                     AgentCard
====================================================
--- General ---
Name        : Hello World Agent - Extended Edition
Description : The full-featured hello world agent for authenticated users.
Version     : 0.0.2

--- Interfaces ---
  [0] http://127.0.0.1:9999  (JSONRPC 1.0)

--- Capabilities ---
Streaming           : True
Push notifications  : False
Extended agent card : True

--- I/O Modes ---
Input  : text/plain
Output : text/plain

--- Skills ---
----------------------------------------------------
  ID          : echo_bot
  Name        : Echo Bot
  Description : An example agent that acknowledges client request and responds with a "Hello World" message.
  Tags        : a2a, echo-example
  Example     : hi
  Example     : how are you
----------------------------------------------------
  ID          : echo_bot_super_mode
  Name        : Echo Bot (Super Mode)
  Description : An extended version of Echo Bot that responds with extra enthusiasm!
  Tags        : a2a, echo-example, extended
  Example     : super hi
  Example     : give me a super hello
====================================================
--
user >

If the server goes down, the client vomits and crashes

(.venv) isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/helloworld$ python test_client.py

Starting an internactive session with A2A Server [http://127.0.0.1:9999]
Use `exit` to quit.
user > hello
Initializes the A2ACardResolver instance with an HTTP client
Traceback (most recent call last):
  File "/home/isaac/Workspaces/a2a-samples/samples/python/agents/helloworld/.venv/lib/python3.14/site-packages/httpx/_transports/default.py", line 101, in map_httpcore_exceptions
    yield
  File "/home/isaac/Workspaces/a2a-samples/samples/python/agents/helloworld/.venv/lib/python3.14/site-packages/httpx/_transports/default.py", line 394, in handle_async_request
    resp = await self._pool.handle_async_request(req)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/isaac/Workspaces/a2a-samples/samples/python/agents/helloworld/.venv/lib/python3.14/site-packages/httpcore/_async/connection_pool.py", line 256, in handle_async_request
    raise exc from None
  File "/home/isaac/Workspaces/a2a-samples/samples/python/agents/helloworld/.venv/lib/python3.14/site-packages/httpcore/_async/connection_pool.py", line 236, in handle_async_request
    response = await connection.handle_async_request(
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
        pool_request.request
        ^^^^^^^^^^^^^^^^^^^^
    )
    ^
  File "/home/isaac/Workspaces/a2a-samples/samples/python/agents/helloworld/.venv/lib/python3.14/site-packages/httpcore/_async/connection.py", line 101, in handle_async_request
    raise exc
  File "/home/isaac/Workspaces/a2a-samples/samples/python/agents/helloworld/.venv/lib/python3.14/site-packages/httpcore/_async/connection.py", line 78, in handle_async_request
    stream = await self._connect(request)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/isaac/Workspaces/a2a-samples/samples/python/agents/helloworld/.venv/lib/python3.14/site-packages/httpcore/_async/connection.py", line 124, in _connect
    stream = await self._network_backend.connect_tcp(**kwargs)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/isaac/Workspaces/a2a-samples/samples/python/agents/helloworld/.venv/lib/python3.14/site-packages/httpcore/_backends/auto.py", line 31, in connect_tcp
    return await self._backend.connect_tcp(
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    ...<5 lines>...
    )
    ^

Adding in Vertex AI (LLM) support

Let’s now look at the adk_expense_reimbursement sample

The interesting parts here are the return form which will return back a dictionary object

def return_form(
    form_request: dict[str, Any],
    tool_context: ToolContext,
    instructions: Optional[str] = None,
) -> dict[str, Any]:
    """Returns a structured json object indicating a form to complete.

    Args:
        form_request (dict[str, Any]): The request form data.
        tool_context (ToolContext): The context in which the tool operates.
        instructions (str): Instructions for processing the form. Can be an empty string.

    Returns:
        dict[str, Any]: A JSON dictionary for the form response.
    """
    if isinstance(form_request, str):
        form_request = json.loads(form_request)

    tool_context.actions.skip_summarization = True
    tool_context.actions.escalate = True
    form_dict = {
        'type': 'form',
        'form': {
            'type': 'object',
            'properties': {
                'date': {
                    'type': 'string',
                    'format': 'date',
                    'description': 'Date of expense',
                    'title': 'Date',
                },
                'amount': {
                    'type': 'string',
                    'format': 'number',
                    'description': 'Amount of expense',
                    'title': 'Amount',
                },
                'purpose': {
                    'type': 'string',
                    'description': 'Purpose of expense',
                    'title': 'Purpose',
                },
                'request_id': {
                    'type': 'string',
                    'description': 'Request id',
                    'title': 'Request ID',
                },
            },
            'required': list(form_request.keys()),
        },
        'form_data': form_request,
        'instructions': instructions,
    }
    return json.dumps(form_dict)

A dictionary object makes it easier for other agentic flows to process and use the output.

The “LLM” parts are defined in the ReimbursementAgent class

class ReimbursementAgent:
    """An agent that handles reimbursement requests."""

    SUPPORTED_CONTENT_TYPES = ['text', 'text/plain']

    def __init__(self):
        self._agent = self._build_agent()
        self._user_id = 'remote_agent'
        self._runner = Runner(
            app_name=self._agent.name,
            agent=self._agent,
            artifact_service=InMemoryArtifactService(),
            session_service=InMemorySessionService(),
            memory_service=InMemoryMemoryService(),
        )

    def get_processing_message(self) -> str:
        return 'Processing the reimbursement request...'

    def _build_agent(self) -> LlmAgent:
        """Builds the LLM agent for the reimbursement agent."""
        LITELLM_MODEL = os.getenv('LITELLM_MODEL', 'gemini/gemini-2.0-flash-001')
        return LlmAgent(
            model=LiteLlm(model=LITELLM_MODEL),
            name='reimbursement_agent',
            description=(
                'This agent handles the reimbursement process for the employees'
                ' given the amount and purpose of the reimbursement.'
            ),
            instruction="""
    You are an agent who handles the reimbursement process for employees.

    When you receive a reimbursement request, you should first create a new request form using create_request_form(). Only provide default values if they are provided by the user, otherwise use an empty string as the default value.
      1. 'Date': the date of the transaction.
      2. 'Amount': the dollar amount of the transaction.
      3. 'Business Justification/Purpose': the reason for the reimbursement.

    Once you created the form, you should return the result of calling return_form with the form data from the create_request_form call.

    Once you received the filled-out form back from the user, you should then check the form contains all required information:
      1. 'Date': the date of the transaction.
      2. 'Amount': the value of the amount of the reimbursement being requested.
      3. 'Business Justification/Purpose': the item/object/artifact of the reimbursement.

    If you don't have all of the information, you should reject the request directly by calling the request_form method, providing the missing fields.


    For valid reimbursement requests, you can then use reimburse() to reimburse the employee.
      * In your response, you should include the request_id and the status of the reimbursement request.

    """,
            tools=[
                create_request_form,
                reimburse,
                return_form,
            ],
        )

For this demo I’ll need a GEMINI API KEY

I’ll create one in the API Keys area of Agent Platform

/img/2026-10-a2a-02.png

I’ll put that key in a .env file (not going to show it here obviously)

isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/adk_expense_reimbursement$ vi .env
isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/adk_expense_reimbursement$ cat .env | sed 's/=.*/=*******************/'
GEMINI_API_KEY=*******************

We can now fire up the agent server with uvicorn

isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/adk_expense_reimbursement$ uv run .
INFO:     Started server process [1399179]
INFO:     Waiting for application startup.
INFO:     Application startup complete.
INFO:     Uvicorn running on http://localhost:10002 (Press CTRL+C to quit)

I can use the Hosts CLI for this test connecting to the port we see above (10002)

(.venv) isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/hosts/cli$ uv run . --agent http://localhost:10002
warning: `VIRTUAL_ENV=/home/isaac/Workspaces/a2a-samples/samples/python/agents/helloworld/.venv` does not match the project environment path `/home/isaac/Workspaces/a2a-samples/samples/python/.venv` and will be ignored; use `--active` to target the active environment instead
Will use headers: {}
======= Agent Card ========
{"capabilities":{"streaming":true},"defaultInputModes":["text","text/plain"],"defaultOutputModes":["text","text/plain"],"description":"This agent handles the reimbursement process for the employees given the amount and purpose of the reimbursement.","name":"Reimbursement Agent","preferredTransport":"JSONRPC","protocolVersion":"0.3.0","skills":[{"description":"Helps with the reimbursement process for users given the amount and purpose of the reimbursement.","examples":["Can you reimburse me $20 for my lunch with the clients?"],"id":"process_reimbursement","name":"Process Reimbursement Tool","tags":["reimbursement"]}],"url":"http://localhost:10002/","version":"1.0.0"}
/home/isaac/Workspaces/a2a-samples/samples/python/hosts/cli/__main__.py:102: DeprecationWarning: A2AClient is deprecated and will be removed in a future version. Use ClientFactory to create a client with a JSON-RPC transport.
  client = A2AClient(httpx_client, agent_card=card)
=========  starting a new task ========

What do you want to send to the agent? (:q or quit to exit):

However, in testing, the mere presence of an API key didn’t work

This is because that Key we created is rather limited by default to just some Agent Platform APIs

Let’s fix that. We’ll click “Edit Key”

/img/2026-10-a2a-04.png

But wait, I don’t even have Generative Language as an option on this project

/img/2026-10-a2a-05.png

I can enable Generative AI on this project

/img/2026-10-a2a-06.png

Danger Notes: Once GenAI is enabled, you have to be very careful with the project as that can skyrocket costs if compromised

/img/2026-10-a2a-07.png

It automatically bound the API key to an SA (vertex-express) and thus I am able to select “Gemini API” now

/img/2026-10-a2a-08.png

Because of my fears of misuse, I’m right away selecting my current egress IP (you can use sites like https://whatismyip.com if don’t know yours) as a restriction

/img/2026-10-a2a-09.png

However, I kept getting an error trying to save

/img/2026-10-a2a-10.png

It might just take a while to take effect. I have a different key that is enabled in this way (with IP restrictions and “Gemini API” access) so I used that in my .env so I could move forward

This time I got an error in the agent (server):

|   File "/home/isaac/Workspaces/a2a-samples/samples/python/.venv/lib/python3.13/site-packages/litellm/litellm_core_utils/exception_mapping_utils.py", line 2301, in exception_type
|     raise e
|   File "/home/isaac/Workspaces/a2a-samples/samples/python/.venv/lib/python3.13/site-packages/litellm/litellm_core_utils/exception_mapping_utils.py", line 1315, in exception_type
|     raise NotFoundError(
|     ...<3 lines>...
|     )
| litellm.exceptions.NotFoundError: litellm.NotFoundError: VertexAIException - {
|   "error": {
|     "code": 404,
|     "message": "This model models/gemini-2.0-flash-001 is no longer available. Please update your code to use models/gemini-3.8-flash for the latest features and improvements. We recommend you to use the Interactions API (https://ai.google.dev/gemini-api/docs/get-started).",
|     "status": "NOT_FOUND"
|   }
| }

As prompted above, I swapped up “2.0” for “3.8” flash

isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/adk_expense_reimbursement$ git diff agent.py
diff --git a/samples/python/agents/adk_expense_reimbursement/agent.py b/samples/python/agents/adk_expense_reimbursement/agent.py
index 4d04245..b246aa0 100644
--- a/samples/python/agents/adk_expense_reimbursement/agent.py
+++ b/samples/python/agents/adk_expense_reimbursement/agent.py
@@ -134,7 +134,7 @@ class ReimbursementAgent:

     def _build_agent(self) -> LlmAgent:
         """Builds the LLM agent for the reimbursement agent."""
-        LITELLM_MODEL = os.getenv('LITELLM_MODEL', 'gemini/gemini-2.0-flash-001')
+        LITELLM_MODEL = os.getenv('LITELLM_MODEL', 'gemini/gemini-3.8-flash')
         return LlmAgent(
             model=LiteLlm(model=LITELLM_MODEL),
             name='reimbursement_agent',

This time it seemed like it might work, but crashed again

| litellm.exceptions.BadRequestError: litellm.BadRequestError: VertexAIException BadRequestError - {
|   "error": {
|     "code": 400,
|     "message": "Function call is missing a thought_signature in functionCall parts. This is required for tools to work correctly, and missing thought_signature may lead to degraded model performance. Additional data, function call `default_api:create_request_form` , position 2. Please refer to https://ai.google.dev/gemini-api/docs/thought-signatures for more details.",
|     "status": "INVALID_ARGUMENT"
|   }

That output links to a page that says how to see our thought tokens, but not pass them

/img/2026-10-a2a-11.png

uv pip install –dry-run -U google-adk

I wish Gemini 3.8 flash in Antigravity would work, but sometimes it’s just really not up to the task. Even feeding it the exact Github issue it could sort it out.

/img/2026-10-a2a-12.png

I tried to search for code examples online, but couldn’t find anything similar.

I gave the same issue to Github Copilot which selected Claude Sonnet for the task (when you leave it as “Auto” you save 10% on whatever model it picks)

/img/2026-10-a2a-13.png

That worked (switching from LiteLLM to Googles library)

isaac@isaac-G707:~/Workspaces/a2a-samples/samples/python/agents/adk_expense_reimbursement$ uv run .
INFO:     Started server process [1451456]
INFO:     Waiting for application startup.
INFO:     Application startup complete.
INFO:     Uvicorn running on http://localhost:10002 (Press CTRL+C to quit)
INFO:     127.0.0.1:50144 - "GET /.well-known/agent-card.json HTTP/1.1" 200 OK
INFO:     127.0.0.1:37662 - "POST / HTTP/1.1" 200 OK
INFO:google_adk.google.adk.models.google_llm:Sending out request, model: gemini-3.8-flash, backend: GoogleLLMVariant.GEMINI_API, stream: False
INFO:google_genai.models:AFC is enabled with max remote calls: 10.
INFO:google_adk.google.adk.models.google_llm:Response received from the model.
WARNING:google_genai.types:Warning: there are non-text parts in the response: ['thought_signature', 'function_call'], returning concatenated text result from text parts. Check the full candidates.content.parts accessor to get the full model response.
INFO:google_adk.google.adk.models.google_llm:Sending out request, model: gemini-3.8-flash, backend: GoogleLLMVariant.GEMINI_API, stream: False
INFO:google_genai.models:AFC is enabled with max remote calls: 10.
INFO:google_adk.google.adk.models.google_llm:Response received from the model.
WARNING:google_genai.types:Warning: there are non-text parts in the response: ['thought_signature', 'function_call'], returning concatenated text result from text parts. Check the full candidates.content.parts accessor to get the full model response.
INFO:     127.0.0.1:37662 - "POST / HTTP/1.1" 200 OK

The output is a bit verbose, but we can see I ask “Can I expense $20 for Stewart”

=========  starting a new task ========

What do you want to send to the agent? (:q or quit to exit): Can I expense $20 for Stewart?
Select a file path to attach? (press enter to skip):
stream event => {"contextId":"bac760febb7d4b349b0523a4ba235d2f","history":[{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"5e9c680a-9cca-47e3-a36b-f70d2bb6247c","parts":[{"kind":"text","text":"Can I expense $20 for Stewart?"}],"role":"user","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"}],"id":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a","kind":"task","status":{"state":"submitted"}}
stream event => {"contextId":"bac760febb7d4b349b0523a4ba235d2f","final":false,"kind":"status-update","status":{"message":{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"f4d343b2-eb59-4e3a-92dd-69f5b6819bca","parts":[{"kind":"text","text":"Processing the reimbursement request..."}],"role":"agent","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"},"state":"working","timestamp":"2026-10-02T12:37:37.141623+00:00"},"taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"}
stream event => {"contextId":"bac760febb7d4b349b0523a4ba235d2f","final":false,"kind":"status-update","status":{"message":{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"c07d8a76-79ea-42df-83f2-3e85d841cac1","parts":[{"kind":"text","text":"Processing the reimbursement request..."}],"role":"agent","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"},"state":"working","timestamp":"2026-10-02T12:37:37.142281+00:00"},"taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"}
stream event => {"contextId":"bac760febb7d4b349b0523a4ba235d2f","final":false,"kind":"status-update","status":{"message":{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"9813eab6-a7f0-4943-aac1-b1770ef5b224","parts":[{"kind":"text","text":"Processing the reimbursement request..."}],"role":"agent","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"},"state":"working","timestamp":"2026-10-02T12:37:38.583959+00:00"},"taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"}
stream event => {"contextId":"bac760febb7d4b349b0523a4ba235d2f","final":true,"kind":"status-update","status":{"message":{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"80f68e8e-8fc3-4cd9-b688-1a524d022266","parts":[{"data":{"type":"form","form":{"type":"object","properties":{"date":{"type":"string","format":"date","description":"Date of expense","title":"Date"},"amount":{"type":"string","format":"number","description":"Amount of expense","title":"Amount"},"purpose":{"type":"string","description":"Purpose of expense","title":"Purpose"},"request_id":{"type":"string","description":"Request id","title":"Request ID"}},"required":["purpose","date","amount","request_id"]},"form_data":{"purpose":"for Stewart","date":"<transaction date>","amount":"20","request_id":"request_id_7470669"},"instructions":"Please provide the transaction date and confirm the details to complete your reimbursement request."},"kind":"data"}],"role":"agent","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"},"state":"input-required","timestamp":"2026-10-02T12:37:38.584597+00:00"},"taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"}

{"contextId":"bac760febb7d4b349b0523a4ba235d2f","history":[{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"5e9c680a-9cca-47e3-a36b-f70d2bb6247c","parts":[{"kind":"text","text":"Can I expense $20 for Stewart?"}],"role":"user","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"},{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"f4d343b2-eb59-4e3a-92dd-69f5b6819bca","parts":[{"kind":"text","text":"Processing the reimbursement request..."}],"role":"agent","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"},{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"c07d8a76-79ea-42df-83f2-3e85d841cac1","parts":[{"kind":"text","text":"Processing the reimbursement request..."}],"role":"agent","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"},{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"9813eab6-a7f0-4943-aac1-b1770ef5b224","parts":[{"kind":"text","text":"Processing the reimbursement request..."}],"role":"agent","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"}],"id":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a","kind":"task","status":{"message":{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"80f68e8e-8fc3-4cd9-b688-1a524d022266","parts":[{"data":{"type":"form","form":{"type":"object","properties":{"date":{"type":"string","format":"date","description":"Date of expense","title":"Date"},"amount":{"type":"string","format":"number","description":"Amount of expense","title":"Amount"},"purpose":{"type":"string","description":"Purpose of expense","title":"Purpose"},"request_id":{"type":"string","description":"Request id","title":"Request ID"}},"required":["purpose","date","amount","request_id"]},"form_data":{"purpose":"for Stewart","date":"<transaction date>","amount":"20","request_id":"request_id_7470669"},"instructions":"Please provide the transaction date and confirm the details to complete your reimbursement request."},"kind":"data"}],"role":"agent","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"},"state":"input-required","timestamp":"2026-10-02T12:37:38.584597+00:00"}}

And in that response it says “Please provide the transaction date and confirm the details to complete your reimbursement request”

I then gave it a date and we can see it got to a point of saying Approved


What do you want to send to the agent? (:q or quit to exit): The transaction date was 2026-08-30
Select a file path to attach? (press enter to skip):
stream event => {"contextId":"bac760febb7d4b349b0523a4ba235d2f","final":false,"kind":"status-update","status":{"message":{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"6208bb89-fd19-46ae-848a-9ae236a4bfed","parts":[{"kind":"text","text":"Processing the reimbursement request..."}],"role":"agent","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"},"state":"working","timestamp":"2026-10-02T12:39:26.254249+00:00"},"taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"}
stream event => {"contextId":"bac760febb7d4b349b0523a4ba235d2f","final":false,"kind":"status-update","status":{"message":{"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"message","messageId":"a4295f92-43b6-4675-944d-f1e34153558c","parts":[{"kind":"text","text":"Processing the reimbursement request..."}],"role":"agent","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"},"state":"working","timestamp":"2026-10-02T12:39:26.254671+00:00"},"taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"}
stream event => {"artifact":{"artifactId":"1c0004cf-4bdb-44e1-8692-39162631a8ca","name":"form","parts":[{"kind":"text","text":"Your reimbursement request has been processed.\n\n* **Request ID:** request_id_7470669\n* **Status:** Approved"}]},"contextId":"bac760febb7d4b349b0523a4ba235d2f","kind":"artifact-update","taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"}
stream event => {"contextId":"bac760febb7d4b349b0523a4ba235d2f","final":true,"kind":"status-update","status":{"state":"completed","timestamp":"2026-10-02T12:39:28.101407+00:00"},"taskId":"eeada2cd-9fd6-4d3f-b5de-33458d3a391a"}
=========  starting a new task ========

Back in Google Cloud I was impressed how quickly I saw some stats on Usage

/img/2026-10-a2a-14.png

Now, let’s do it differently

Let’s start with a RESTful service that could do hello world using Python’s FastAPI framework. I let Agy pick a decent TUI for the caller

/img/2026-10-a2a-15.png

Gemini 3.8 Flash yet again really disappointed me. This was a super straightforward request

/img/2026-10-a2a-16.png

I pivoted to 3.1 Pro to see if that would do better

/img/2026-10-a2a-17.png

It created exactly what I was thinking

/img/2026-10-a2a-18.png

And with far less time and tokens

/img/2026-10-a2a-19.png

I’ll now do the setup

isaac@isaac-G707:~/Workspaces/do_it_different$ python3 -m venv venv
isaac@isaac-G707:~/Workspaces/do_it_different$ source venv/bin/activate                                                         (venv) isaac@isaac-G707:~/Workspaces/do_it_different$ pip install -r requirements.txt
Collecting fastapi (from -r requirements.txt (line 1))
  Downloading fastapi-0.142.2-py3-none-any.whl.metadata (27 kB)
Collecting requests (from -r requirements.txt (line 3))
  Using cached requests-2.34.2-py3-none-any.whl.metadata (4.8 kB)                                                               Collecting pydantic (from -r requirements.txt (line 4))
  Using cached pydantic-2.13.5-py3-none-any.whl.metadata (110 kB)
Collecting uvicorn[standard] (from -r requirements.txt (line 2))                                                                  Using cached uvicorn-0.54.0-py3-none-any.whl.metadata (6.6 kB)
Collecting starlette>=0.46.0 (from fastapi->-r requirements.txt (line 1))
  Using cached starlette-1.7.0-py3-none-any.whl.metadata (6.6 kB)
Collecting typing-extensions>=4.8.0 (from fastapi->-r requirements.txt (line 1))                                                  Using cached typing_extensions-4.16.0-py3-none-any.whl.metadata (3.3 kB)
Collecting typing-inspection>=0.4.2 (from fastapi->-r requirements.txt (line 1))
  Using cached typing_inspection-0.4.4-py3-none-any.whl.metadata (2.6 kB)
Collecting annotated-doc>=0.0.2 (from fastapi->-r requirements.txt (line 1))                                                      Using cached annotated_doc-0.0.5-py3-none-any.whl.metadata (6.5 kB)
Collecting opentelemetry-api>=1.44.0 (from fastapi->-r requirements.txt (line 1))
  Using cached opentelemetry_api-1.45.0-py3-none-any.whl.metadata (1.4 kB)
Collecting click>=7.0 (from uvicorn[standard]->-r requirements.txt (line 2))
  Using cached click-8.5.0-py3-none-any.whl.metadata (2.6 kB)
Collecting h11>=0.8 (from uvicorn[standard]->-r requirements.txt (line 2))
  Using cached h11-0.16.0-py3-none-any.whl.metadata (8.3 kB)                                                                    Collecting httptools>=0.8.0 (from uvicorn[standard]->-r requirements.txt (line 2))
  Using cached httptools-0.8.0-cp314-cp314-manylinux1_x86_64.manylinux_2_28_x86_64.manylinux_2_5_x86_64.whl.metadata (3.5 kB)
Collecting python-dotenv>=0.13 (from uvicorn[standard]->-r requirements.txt (line 2))                                             Downloading python_dotenv-1.2.4-py3-none-any.whl.metadata (29 kB)
Collecting pyyaml>=5.1 (from uvicorn[standard]->-r requirements.txt (line 2))
  Using cached pyyaml-6.0.3-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl.metadata (2.4 kB)
Collecting uvloop>=0.15.1 (from uvicorn[standard]->-r requirements.txt (line 2))                                                  Downloading uvloop-0.23.0-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl.metadata (5.1 kB)
Collecting watchfiles>=0.20 (from uvicorn[standard]->-r requirements.txt (line 2))
  Downloading watchfiles-1.3.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (4.9 kB)
Collecting websockets>=13.0 (from uvicorn[standard]->-r requirements.txt (line 2))
  Using cached websockets-17.1-cp314-cp314-manylinux1_x86_64.manylinux_2_28_x86_64.manylinux_2_5_x86_64.whl.metadata (6.3 kB)
Collecting charset_normalizer<4,>=2 (from requests->-r requirements.txt (line 3))
  Downloading charset_normalizer-3.5.2-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl.metadata (46 kB)
Collecting idna<4,>=2.5 (from requests->-r requirements.txt (line 3))
  Using cached idna-3.20-py3-none-any.whl.metadata (7.2 kB)
Collecting urllib3<3,>=1.26 (from requests->-r requirements.txt (line 3))                                                         Using cached urllib3-2.8.0-py3-none-any.whl.metadata (7.4 kB)
Collecting certifi>=2023.5.7 (from requests->-r requirements.txt (line 3))
  Using cached certifi-2026.7.22-py3-none-any.whl.metadata (2.5 kB)                                                             Collecting annotated-types>=0.6.0 (from pydantic->-r requirements.txt (line 4))
  Using cached annotated_types-0.8.0-py3-none-any.whl.metadata (15 kB)
Collecting pydantic-core==2.46.5 (from pydantic->-r requirements.txt (line 4))
  Using cached pydantic_core-2.46.5-cp314-cp314-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (6.6 kB)
Collecting anyio<5,>=4.0.0 (from starlette>=0.46.0->fastapi->-r requirements.txt (line 1))
  Using cached anyio-4.15.1-py3-none-any.whl.metadata (4.7 kB)
Downloading fastapi-0.142.2-py3-none-any.whl (144 kB)
Using cached uvicorn-0.54.0-py3-none-any.whl (87 kB)                                                                            Using cached requests-2.34.2-py3-none-any.whl (73 kB)                                                                           Downloading charset_normalizer-3.5.2-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl (255 kB)
Using cached idna-3.20-py3-none-any.whl (69 kB)
Using cached urllib3-2.8.0-py3-none-any.whl (135 kB)
Using cached pydantic-2.13.5-py3-none-any.whl (472 kB)
Using cached pydantic_core-2.46.5-cp314-cp314-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.1 MB)
Using cached annotated_doc-0.0.5-py3-none-any.whl (5.3 kB)
Using cached annotated_types-0.8.0-py3-none-any.whl (13 kB)
Using cached certifi-2026.7.22-py3-none-any.whl (136 kB)
Using cached click-8.5.0-py3-none-any.whl (125 kB)
Using cached h11-0.16.0-py3-none-any.whl (37 kB)
Using cached httptools-0.8.0-cp314-cp314-manylinux1_x86_64.manylinux_2_28_x86_64.manylinux_2_5_x86_64.whl (481 kB)
Using cached opentelemetry_api-1.45.0-py3-none-any.whl (60 kB)
Downloading python_dotenv-1.2.4-py3-none-any.whl (23 kB)
Using cached pyyaml-6.0.3-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl (794 kB)
Using cached starlette-1.7.0-py3-none-any.whl (78 kB)
Using cached anyio-4.15.1-py3-none-any.whl (132 kB)
Using cached typing_extensions-4.16.0-py3-none-any.whl (45 kB)
                                                                  Using cached typing_inspection-0.4.4-py3-none-any.whl (14 kB)
Downloading uvloop-0.23.0-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl (4.4 MB)
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 4.4/4.4 MB 12.0 MB/s eta 0:00:00                                                    Downloading watchfiles-1.3.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (458 kB)
Using cached websockets-17.1-cp314-cp314-manylinux1_x86_64.manylinux_2_28_x86_64.manylinux_2_5_x86_64.whl (224 kB)
Installing collected packages: websockets, uvloop, urllib3, typing-extensions, pyyaml, python-dotenv, idna, httptools, h11, click, charset_normalizer, certifi, annotated-types, annotated-doc, uvicorn, typing-inspection, requests, pydantic-core, opentelemetry-api, anyio, watchfiles, starlette, pydantic, fastapi
Successfully installed annotated-doc-0.0.5 annotated-types-0.8.0 anyio-4.15.1 certifi-2026.7.22 charset_normalizer-3.5.2 click-8.5.0 fastapi-0.142.2 h11-0.16.0 httptools-0.8.0 idna-3.20 opentelemetry-api-1.45.0 pydantic-2.13.5 pydantic-core-2.46.5 python-dotenv-1.2.4 pyyaml-6.0.3 requests-2.34.2 starlette-1.7.0 typing-extensions-4.16.0 typing-inspection-0.4.4 urllib3-2.8.0 uvicorn-0.54.0 uvloop-0.23.0 watchfiles-1.3.0 websockets-17.1

And fire up the server

(venv) isaac@isaac-G707:~/Workspaces/do_it_different$ uvicorn server:app --reload
INFO:     Will watch for changes in these directories: ['/home/isaac/Workspaces/do_it_different']
INFO:     Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)                                                       INFO:     Started reloader process [1475222] using WatchFiles
INFO:     Started server process [1475224]
INFO:     Waiting for application startup.                                                                                      INFO:     Application startup complete.

This was faster, lighter and more performant

The requirements.txt:

$ cat requirements.txt
fastapi
uvicorn[standard]
requests
pydantic

The cli.py

$ cat cli.py
import sys
import requests

def main():
    print("Welcome to the Echo Bot CLI!")
    print("Type your message and press Enter. Type 'quit' or 'exit' to stop.")

    url = "http://127.0.0.1:8000/echo_bot"

    while True:
        try:
            user_input = input("\nEnter text: ")
        except (KeyboardInterrupt, EOFError):
            print("\nExiting...")
            break

        if user_input.strip().lower() in ['quit', 'exit']:
            print("Exiting...")
            break

        payload = {"user_request": user_input}

        try:
            response = requests.post(url, json=payload)
            response.raise_for_status()
            data = response.json()
            print(f"Server response: {data.get('response')}")
        except requests.exceptions.ConnectionError:
            print("Error: Could not connect to the server at http://127.0.0.1:8000")
            print("Please ensure the FastAPI server is running (e.g. 'uvicorn server:app').")
        except Exception as e:
            print(f"An error occurred: {e}")

if __name__ == "__main__":
    main()

And lastly the server python

$ cat server.py
from fastapi import FastAPI, Query
from pydantic import BaseModel
from typing import Optional

app = FastAPI(title="Echo Bot API")

class EchoRequest(BaseModel):
    user_request: Optional[str] = ""

def process_request(user_request: Optional[str]) -> dict:
    if not user_request or not user_request.strip():
        return {"response": "No text input is provided!"}

    # Returning the exact string format requested
    return {"response": f"Hello, World! I have received your request ({user_request})"}

@app.post("/echo_bot")
async def echo_bot_post(request: EchoRequest):
    return process_request(request.user_request)

@app.get("/echo_bot")
async def echo_bot_get(user_request: Optional[str] = Query("")):
    return process_request(user_request)

if __name__ == "__main__":
    import uvicorn
    uvicorn.run("server:app", host="127.0.0.1", port=8000, reload=True)

Okay, perhaps a hello world example doesn’t really show the value of A2A

Expense agent

I’ll use Claude Sonnet 5.5 via Copilot on this ask

/img/2026-10-a2a-21.png

Similar to before, I’ll setup the environment

isaac@isaac-G707:~/Workspaces/do_it_different/expense_reimbursement$ python3 -m venv venv
isaac@isaac-G707:~/Workspaces/do_it_different/expense_reimbursement$ source venv/bin/activate
(venv) isaac@isaac-G707:~/Workspaces/do_it_different/expense_reimbursement$ pip install -r requirements.txt
Collecting fastapi (from -r requirements.txt (line 1))
  Using cached fastapi-0.142.2-py3-none-any.whl.metadata (27 kB)
Collecting uvicorn (from -r requirements.txt (line 2))
  Using cached uvicorn-0.54.0-py3-none-any.whl.metadata (6.6 kB)
Collecting litellm (from -r requirements.txt (line 3))
  Downloading litellm-1.103.2-cp310-abi3-manylinux_2_28_x86_64.whl.metadata (42 kB)
Collecting google-cloud-aiplatform (from -r requirements.txt (line 4))
  Downloading google_cloud_aiplatform-2.3.0-py2.py3-none-any.whl.metadata (52 kB)
Collecting requests (from -r requirements.txt (line 5))
  Using cached requests-2.34.2-py3-none-any.whl.metadata (4.8 kB)
Collecting starlette>=0.46.0 (from fastapi->-r requirements.txt (line 1))
  Using cached starlette-1.7.0-py3-none-any.whl.metadata (6.6 kB)
Collecting pydantic>=2.9.0 (from fastapi->-r requirements.txt (line 1))
  Using cached pydantic-2.13.5-py3-none-any.whl.metadata (110 kB)
Collecting typing-extensions>=4.8.0 (from fastapi->-r requirements.txt (line 1))
  Using cached typing_extensions-4.16.0-py3-none-any.whl.metadata (3.3 kB)
Collecting typing-inspection>=0.4.2 (from fastapi->-r requirements.txt (line 1))
  Using cached typing_inspection-0.4.4-py3-none-any.whl.metadata (2.6 kB)
Collecting annotated-doc>=0.0.2 (from fastapi->-r requirements.txt (line 1))
  Using cached annotated_doc-0.0.5-py3-none-any.whl.metadata (6.5 kB)
Collecting opentelemetry-api>=1.44.0 (from fastapi->-r requirements.txt (line 1))
  Using cached opentelemetry_api-1.45.0-py3-none-any.whl.metadata (1.4 kB)
Collecting click>=7.0 (from uvicorn->-r requirements.txt (line 2))
  Using cached click-8.5.0-py3-none-any.whl.metadata (2.6 kB)
Collecting h11>=0.8 (from uvicorn->-r requirements.txt (line 2))
  Using cached h11-0.16.0-py3-none-any.whl.metadata (8.3 kB)
Collecting fastuuid<1.0,>=0.14.0 (from litellm->-r requirements.txt (line 3))
  Downloading fastuuid-0.14.0-cp314-cp314-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (1.1 kB)
Collecting httpx<1.0,>=0.28.0 (from httpx[http2]<1.0,>=0.28.0->litellm->-r requirements.txt (line 3))
  Using cached httpx-0.28.1-py3-none-any.whl.metadata (7.1 kB)

... snip ...

I needed to login with this implementation with application creds

$ gcloud auth application-default login
Your browser has been opened to visit:

    https://accounts.google.com/o/o.....

Gtk-Message: 08:43:14.376: Not loading module "atk-bridge": The functionality is provided by GTK natively. Please try to not load it.

Credentials saved to file: [/home/isaac/.config/gcloud/application_default_credentials.json]

These credentials will be used by any library that requests Application Default Credentials (ADC).
API [cloudresourcemanager.googleapis.com] not enabled on project [careful-compass-241122]. Would you like to enable and retry
(this will take a few minutes)? (y/N)?  y

Enabling service [cloudresourcemanager.googleapis.com] on project [careful-compass-241122]...
Operation "operations/acat.p2-243917726587-45e66113-7f54-4d65-b8d7-fd29a59a2ab3" finished successfully.

Quota project "careful-compass-241122" was added to ADC which can be used by Google client libraries for billing and quota. Note that some services may still bill the project owning the resource.

I’ll try now firing up the server and making sure to also use the latest 3.8 flash model

$ VERTEXAI_PROJECT=careful-compass-241122 LITELLM_MODEL=vertex_ai/gemini-3.8-flash VERTEXAI_LOCATION=global python server.py
INFO:     Will watch for changes in these directories: ['/home/isaac/Workspaces/do_it_different/expense_reimbursement']
INFO:     Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)
INFO:     Started reloader process [1479534] using StatReload
INFO:     Started server process [1479542]
INFO:     Waiting for application startup.
INFO:     Application startup complete.

Like before it detected I was missing a date field, albeit in my reply i didn’t realize i was confirming the values it found so i put the date in for the name (and transaction date), but it worked

$ python cli.py
Expense Reimbursement CLI. Type 'quit' or 'exit' to stop.

You: Can i expense $30 for David?

Please complete the form:
Instructions: Please provide the date of the transaction to complete your reimbursement request.
purpose [David]: 2026-09-05
amount [30]:
date []: 2026-09-05

Agent: Your reimbursement request has been processed.

- **Request ID:** request_id_4818776
- **Status:** approved

You:

As you can see, this worked just fine

As you can see in the agent code, as before we are invoking with LiteLLM and keeping the session in memory as we go

$ cat agent.py
import json
import os
import random
from typing import Any, Optional

import litellm

MODEL = os.getenv("LITELLM_MODEL", "vertex_ai/gemini-2.5-flash")
MAX_TOOL_ROUNDS = 8

SYSTEM_PROMPT = """
You are an agent who handles the reimbursement process for employees.

When you receive a reimbursement request, first create a new request form using create_request_form(). Only provide values that the user supplied, otherwise use an empty string.
  1. 'date': the date of the transaction.
  2. 'amount': the dollar amount of the transaction.
  3. 'purpose': the business justification/purpose of the reimbursement.

Once you created the form, return it to the user by calling return_form with the form data from the create_request_form call.

Once you receive the filled-out form back from the user, check it contains all required information (date, amount, purpose).
If any is missing, call return_form again with the missing fields.

For valid reimbursement requests, use reimburse() to reimburse the employee.
In your response, include the request_id and the status of the reimbursement request.
"""

# Request ids created in this process (demo only).
request_ids: set[str] = set()


def create_request_form(
    date: Optional[str] = None,
    amount: Optional[str] = None,
    purpose: Optional[str] = None,
) -> dict[str, Any]:
    request_id = "request_id_" + str(random.randint(1000000, 9999999))
    request_ids.add(request_id)
    return {
        "request_id": request_id,
        "date": date or "<transaction date>",
        "amount": amount or "<transaction dollar amount>",
        "purpose": purpose or "<business justification/purpose of the transaction>",
    }


def return_form(form_request: Any, instructions: Optional[str] = None) -> dict[str, Any]:
    if isinstance(form_request, str):
        form_request = json.loads(form_request)
    return {
        "type": "form",
        "form": {
            "type": "object",
            "properties": {
                "date": {"type": "string", "format": "date", "title": "Date"},
                "amount": {"type": "string", "format": "number", "title": "Amount"},
                "purpose": {"type": "string", "title": "Purpose"},
                "request_id": {"type": "string", "title": "Request ID"},
            },
            "required": list(form_request.keys()),
        },
        "form_data": form_request,
        "instructions": instructions,
    }


def reimburse(request_id: str) -> dict[str, Any]:
    if request_id not in request_ids:
        return {"request_id": request_id, "status": "Error: Invalid request_id."}
    return {"request_id": request_id, "status": "approved"}


TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "create_request_form",
            "description": "Create a request form for the employee to fill out. Use empty strings for unknown values.",
            "parameters": {
                "type": "object",
                "properties": {
                    "date": {"type": "string", "description": "Date of the transaction."},
                    "amount": {"type": "string", "description": "Requested amount."},
                    "purpose": {"type": "string", "description": "Purpose of the request."},
                },
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "return_form",
            "description": "Return a form to the user to complete.",
            "parameters": {
                "type": "object",
                "properties": {
                    "form_request": {
                        "type": "object",
                        "description": "The form data from create_request_form.",
                    },
                    "instructions": {
                        "type": "string",
                        "description": "Instructions for completing the form.",
                    },
                },
                "required": ["form_request"],
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "reimburse",
            "description": "Reimburse the employee for a given request_id.",
            "parameters": {
                "type": "object",
                "properties": {"request_id": {"type": "string"}},
                "required": ["request_id"],
            },
        },
    },
]

TOOL_FUNCS = {
    "create_request_form": create_request_form,
    "return_form": return_form,
    "reimburse": reimburse,
}


class ReimbursementAgent:
    """Keeps per-session message history and runs the LiteLLM tool-calling loop."""

    def __init__(self) -> None:
        self._sessions: dict[str, list[dict[str, Any]]] = {}

    def chat(self, session_id: str, message: str) -> dict[str, Any]:
        messages = self._sessions.setdefault(
            session_id, [{"role": "system", "content": SYSTEM_PROMPT}]
        )
        messages.append({"role": "user", "content": message})

        for _ in range(MAX_TOOL_ROUNDS):
            resp = litellm.completion(model=MODEL, messages=messages, tools=TOOLS)
            msg = resp.choices[0].message
            messages.append(msg.model_dump(exclude_none=True))

            if not msg.tool_calls:
                return {"type": "text", "response": msg.content or ""}

            form = None
            for call in msg.tool_calls:
                result = self._run_tool(call.function.name, call.function.arguments)
                if call.function.name == "return_form" and "error" not in result:
                    form = result
                messages.append(
                    {
                        "role": "tool",
                        "tool_call_id": call.id,
                        "name": call.function.name,
                        "content": json.dumps(result),
                    }
                )
            if form is not None:
                return form

        return {"type": "text", "response": "Error: too many tool-call rounds."}

    @staticmethod
    def _run_tool(name: str, raw_args: str) -> dict[str, Any]:
        func = TOOL_FUNCS.get(name)
        if func is None:
            return {"error": f"Unknown tool: {name}"}
        try:
            return func(**json.loads(raw_args or "{}"))
        except Exception as e:
            return {"error": str(e)}

The server code (similar to the A2A __main__.py) just calls the agent

$ cat server.py
import uuid
from typing import Any, Optional

from fastapi import FastAPI, HTTPException
from pydantic import BaseModel

from agent import ReimbursementAgent

app = FastAPI(title="Expense Reimbursement API")
agent = ReimbursementAgent()


class ChatRequest(BaseModel):
    message: str
    session_id: Optional[str] = None


@app.post("/chat")
def chat(request: ChatRequest) -> dict[str, Any]:
    if not request.message.strip():
        raise HTTPException(status_code=400, detail="message must not be empty")
    session_id = request.session_id or str(uuid.uuid4())
    try:
        result = agent.chat(session_id, request.message)
    except Exception as e:
        raise HTTPException(status_code=502, detail=f"Model call failed: {e}")
    return {"session_id": session_id, **result}


if __name__ == "__main__":
    import uvicorn

    uvicorn.run("server:app", host="127.0.0.1", port=8000, reload=True)

The CLI is just engaging with server

$ cat cli.py
import json
import os

import requests

URL = os.getenv("REIMBURSE_URL", "http://127.0.0.1:8000/chat")


def send(message: str, session_id):
    resp = requests.post(
        URL, json={"message": message, "session_id": session_id}, timeout=120
    )
    resp.raise_for_status()
    return resp.json()


def fill_form(data: dict) -> str:
    if data.get("instructions"):
        print(f"Instructions: {data['instructions']}")
    values = {}
    for key, current in data["form_data"].items():
        if key == "request_id":
            values[key] = current
            continue
        shown = "" if str(current).startswith("<") else current
        entered = input(f"{key} [{shown}]: ").strip()
        values[key] = entered or shown
    return json.dumps(values)


def main():
    print("Expense Reimbursement CLI. Type 'quit' or 'exit' to stop.")
    session_id = None
    while True:
        try:
            message = input("\nYou: ")
        except (KeyboardInterrupt, EOFError):
            print("\nExiting...")
            break
        if message.strip().lower() in ("quit", "exit"):
            break
        if not message.strip():
            continue

        try:
            data = send(message, session_id)
            session_id = data["session_id"]
            while data["type"] == "form":
                print("\nPlease complete the form:")
                data = send(fill_form(data), session_id)
            print(f"\nAgent: {data['response']}")
        except requests.exceptions.ConnectionError:
            print(f"Error: could not connect to {URL}. Is the server running?")
        except Exception as e:
            print(f"An error occurred: {e}")


if __name__ == "__main__":
    main()

And we have a straightforward requirements.txt that includes the google-cloud-aiplatform library

$ cat requirements.txt
fastapi
uvicorn
litellm
google-cloud-aiplatform
requests

However, unlike the A2A implementation, with FastAPI I get OpenAPI endpoints with functional Swagger.

So I can just as easily use the Web UI in Swagger to engage with the agent directly

/img/2026-10-a2a-23.png

Here is a basic Flask app we could use

$ cat webapp.py
import os
import json
import requests
from flask import Flask, render_template_string, request, jsonify

app = Flask(__name__)
URL = os.getenv("REIMBURSE_URL", "http://127.0.0.1:8000/chat")

HTML = """
<!DOCTYPE html>
<html>
<head>
    <title>Expense Reimbursement</title>
    <style>
        body { font-family: sans-serif; max-width: 600px; margin: auto; padding: 20px; }
        #chatbox { height: 400px; border: 1px solid #ccc; overflow-y: scroll; padding: 10px; margin-bottom: 10px; background-color: #fafafa; }
        .msg-container { margin-bottom: 10px; clear: both; overflow: hidden; }
        .user-msg { color: #fff; background-color: #007bff; float: right; padding: 8px 12px; border-radius: 15px; max-width: 80%; }
        .agent-msg { color: #000; background-color: #e9ecef; float: left; padding: 8px 12px; border-radius: 15px; max-width: 80%; }
        .system-msg { color: red; text-align: center; clear: both; }
        #input-area { display: flex; }
        #input-area input[type="text"] { flex: 1; padding: 10px; border: 1px solid #ccc; border-radius: 4px; }
        #input-area button { padding: 10px 15px; margin-left: 5px; cursor: pointer; border: none; background-color: #007bff; color: white; border-radius: 4px; }
        .form-container { background: #fff; padding: 15px; margin-top: 10px; border: 1px solid #ddd; border-radius: 5px; clear: both; float: left; width: 100%; box-sizing: border-box; }
        .form-container label { display: block; margin-top: 10px; font-weight: bold; }
        .form-container input { width: 100%; padding: 8px; margin-top: 5px; box-sizing: border-box; border: 1px solid #ccc; border-radius: 4px; }
        .form-btn { margin-top: 15px; padding: 10px 15px; cursor: pointer; border: none; background-color: #28a745; color: white; border-radius: 4px; width: 100%; }
    </style>
</head>
<body>
    <h2>Expense Reimbursement</h2>
    <div id="chatbox">
        <div class="msg-container">
            <div class="agent-msg">Hello! I am the expense reimbursement agent. What expense report do you have a question about or need to submit?</div>
        </div>
    </div>

    <div id="input-area">
        <input type="text" id="user-input" placeholder="Type your message here..." onkeypress="handleKeyPress(event)">
        <button onclick="sendMessage()">Send</button>
    </div>

    <div id="form-area" style="display: none;">
        <div class="form-container" id="dynamic-form">
            <!-- Form fields go here -->
        </div>
        <button class="form-btn" onclick="submitForm()">Submit Form</button>
    </div>

    <script>
        const chatbox = document.getElementById('chatbox');
        const userInput = document.getElementById('user-input');
        const inputArea = document.getElementById('input-area');
        const formArea = document.getElementById('form-area');
        const dynamicForm = document.getElementById('dynamic-form');
        let currentSessionId = null;

        function appendMessage(sender, text, isHtml=false) {
            const containerDiv = document.createElement('div');
            containerDiv.className = 'msg-container';

            const msgDiv = document.createElement('div');
            if (sender === 'You') {
                msgDiv.className = 'user-msg';
            } else if (sender === 'System') {
                msgDiv.className = 'system-msg';
            } else {
                msgDiv.className = 'agent-msg';
            }

            if(isHtml) {
                msgDiv.innerHTML = text;
            } else {
                msgDiv.innerText = text;
            }

            containerDiv.appendChild(msgDiv);
            chatbox.appendChild(containerDiv);
            chatbox.scrollTop = chatbox.scrollHeight;
        }

        function handleKeyPress(e) {
            if (e.key === 'Enter') {
                sendMessage();
            }
        }

        async function sendToServer(message) {
            try {
                appendMessage('System', 'Loading...', true);
                const loadingMsg = chatbox.lastChild;

                const response = await fetch('/api/chat', {
                    method: 'POST',
                    headers: { 'Content-Type': 'application/json' },
                    body: JSON.stringify({ message: message, session_id: currentSessionId })
                });

                chatbox.removeChild(loadingMsg);

                const data = await response.json();
                if (data.error) {
                    appendMessage('System', 'Error: ' + data.error);
                    inputArea.style.display = 'flex';
                    return;
                }

                currentSessionId = data.session_id;

                if (data.type === 'form') {
                    renderForm(data);
                } else {
                    appendMessage('Agent', data.response);
                    inputArea.style.display = 'flex';
                    formArea.style.display = 'none';
                    userInput.focus();
                }
            } catch (error) {
                appendMessage('System', 'Failed to communicate with server.');
                inputArea.style.display = 'flex';
            }
        }

        function sendMessage() {
            const text = userInput.value.trim();
            if (!text) return;
            appendMessage('You', text);
            userInput.value = '';
            inputArea.style.display = 'none';

            sendToServer(text);
        }

        function renderForm(data) {
            inputArea.style.display = 'none';
            formArea.style.display = 'block';
            dynamicForm.innerHTML = '';

            if (data.instructions) {
                const instDiv = document.createElement('div');
                instDiv.innerText = data.instructions;
                instDiv.style.marginBottom = '10px';
                instDiv.style.color = '#333';
                dynamicForm.appendChild(instDiv);
            }

            for (const [key, value] of Object.entries(data.form_data)) {
                if (key === 'request_id') {
                    const hiddenInput = document.createElement('input');
                    hiddenInput.type = 'hidden';
                    hiddenInput.id = `form-${key}`;
                    hiddenInput.value = value;
                    dynamicForm.appendChild(hiddenInput);
                    continue;
                }

                const label = document.createElement('label');
                label.innerText = key;

                const input = document.createElement('input');
                input.type = 'text';
                input.id = `form-${key}`;

                const shown = String(value).startsWith('<') ? '' : value;
                input.value = shown;
                input.placeholder = String(value).startsWith('<') ? value : '';

                dynamicForm.appendChild(label);
                dynamicForm.appendChild(input);
            }
            chatbox.scrollTop = chatbox.scrollHeight;
        }

        function submitForm() {
            const formData = {};
            const inputs = dynamicForm.querySelectorAll('input');
            inputs.forEach(input => {
                const key = input.id.replace('form-', '');
                formData[key] = input.value;
            });

            appendMessage('You', '<i>[Submitted Form]</i>', true);
            formArea.style.display = 'none';
            sendToServer(JSON.stringify(formData));
        }
    </script>
</body>
</html>
"""

@app.route("/")
def index():
    return render_template_string(HTML)

@app.route("/api/chat", methods=["POST"])
def chat():
    data = request.json
    message = data.get("message", "")
    session_id = data.get("session_id")

    try:
        resp = requests.post(
            URL, json={"message": message, "session_id": session_id}, timeout=120
        )
        resp.raise_for_status()
        return jsonify(resp.json())
    except requests.exceptions.ConnectionError:
        return jsonify({"error": f"could not connect to {URL}. Is the backend server running?"})
    except Exception as e:
        return jsonify({"error": str(e)})

if __name__ == "__main__":
    app.run(port=5000, debug=True)

And in action:

For reference, you can get both the hello world and expense report framework from Github here

Summary

We explored two examples of A2A code. First was a hello world client and server that would just echo back what we send which was easy to fire up. We then explored one tied to AI models that would evaluate an expense report request and format the data.

The session memory seemed well handled in the A2A setup, but I found no benefit over a RESTful service using FastAPI. Moreover, a FastAPI implementation gave us Swagger (OpenAPI) endpoints making it very easy to test use the /docs endpoint and build a nice web interface.

I know I should like A2A but really I see no value in it thus far. With swagger-backed REST services I can easily add head authentication and /docs and /redocs endpoints.

I may come back later and see if there is some benefit in an Agent Registry or some kind of swarm mode, but as for today I see no reason to use A2A for any of my workloads.