This directory contains practical examples for using the infer CLI tool to interact with the Inference Gateway.
- Configure the Inference Gateway server:
# Configure the providers you want to work with and some of the agents credentials (Google Calendar, Context7 - if applicable)
cp .env.example .env- Bring all the containers up:
docker compose up -d- Log the containers verify that everything is up and running:
docker compose ps
docker compose logs -fSet up your CLI configuration via environment variables (review docker-compose.yaml cli service):
INFER_GATEWAY_URL: http://inference-gateway:8080
INFER_A2A_ENABLED: true
INFER_TOOLS_ENABLED: false
INFER_AGENT_MODEL: deepseek/deepseek-v4-pro # Choose whatever LLM you would like to use from the configured providers** Using INFER_A2A_ENABLED: true automatically enables A2A tools (QueryAgent, QueryTask, SubmitTask) even when local tools
are disabled. This simplified configuration gives you only the A2A functionality without needing to configure each
tool individually.
Now you can enter the Interactive Chat within the cli container and start chatting:
docker compose run --rm cliThe browser-agent can run in headed mode with VNC support for real-time viewing of browser automation.
Requirements:
- Browser-agent v0.4.0+ with Xvfb support
BROWSER_HEADLESS: falseenvironment variable set- Shared X11 socket volume between browser-agent and browser-vnc
Usage:
Connect to the VNC server on port 5900:
# Using a VNC client (e.g., RealVNC, TigerVNC, or macOS Screen Sharing)
# Connect to: localhost:5900
# Password: passwordOn macOS, you can use the built-in Screen Sharing app:
open vnc://localhost:5900
# When prompted, enter password: passwordNote: The X11 display is created by Xvfb when the browser-agent container starts. The VNC server will connect automatically and you can view browser automation in real-time.
When the model delegates a task to a configured agent via A2A_SubmitTask, a
live sticky status bar appears just above the input box showing the remote
task's progress:
… conversation viewport …
─────────────────────────────────────────────
◓ Agent(google-calendar-agent=working...)
> _ (input)
When the remote agent completes, the line expands to a tree-style block
showing the token usage and execution statistics from the agent's
Task.metadata (populated by ADK ≥ 0.19.0):
✓ Agent(google-calendar-agent=completed)
├── usage={"prompt_tokens":156,"completion_tokens":89,"total_tokens":245}
└── execution_stats={"iterations":2,"messages":4,"tool_calls":1,"failed_tools":0}
The indicator auto-disappears 5 seconds later so the conversation stays tidy.
Failed tasks show ✗ Agent(<name>=failed: <error>) and follow the same 5s
auto-removal.
Try it: ask the CLI a question that requires the calendar agent (e.g. "What's on my calendar today?") and watch the indicator appear and then auto-clear after the agent responds.
Note: The
usage=…suffix only appears when the remote agent runs ADK ≥ 0.19.0 withEnableUsageMetadataenabled (the default in 0.19.0+). If you seeAgent(name=completed)with nousage=…, the agent image needs to be rebuilt against the newer ADK. The friendly agent name (e.g.mock-agentvs the raw URL) is resolved from~/.infer/agents.yaml- if you've only registered agents viaINFER_A2A_AGENTS, the indicator will show the URL instead.
docker compose run --rm a2a-debugger tasks list
docker compose run --rm a2a-debugger tasks get <task_id>
docker compose run --rm a2a-debugger tasks submit-streaming "What's on my calendar today?"