AppAgent is the core execution runtime in UFO, responsible for carrying out individual subtasks within a specific Windows application. Each AppAgent functions as an isolated, application-specialized worker process launched and orchestrated by the central HostAgent.
 AppAgent Architecture: Application-specialized worker process for subtask execution
AppAgent operates as a child agent under the HostAgent's orchestration:
- Isolated Runtime: Each AppAgent is dedicated to a single Windows application
- Subtask Executor: Executes specific subtasks delegated by HostAgent
- Application Expert: Tailored with deep knowledge of the target app's API surface, control semantics, and domain logic
- Hybrid Execution: Leverages both GUI automation and API-based actions through MCP commands
Unlike monolithic Computer-Using Agents (CUAs) that treat all GUI contexts uniformly, each AppAgent is tailored to a single application and operates with specialized knowledge of its interface and capabilities.
graph TB
subgraph "AppAgent Core Responsibilities"
SR[Sense:<br/>Capture Application State]
RE[Reason:<br/>Analyze Next Action]
EX[Execute:<br/>GUI or API Action]
RP[Report:<br/>Write Results to Blackboard]
end
SR --> RE
RE --> EX
EX --> RP
RP --> SR
style SR fill:#e3f2fd
style RE fill:#fff3e0
style EX fill:#f1f8e9
style RP fill:#fce4ec
| Responsibility | Description | Example |
|---|---|---|
| State Sensing | Capture application UI, detect controls, understand current state | Screenshot Word window → Detect 50 controls → Annotate UI elements |
| Reasoning | Analyze state and determine next action using LLM | "Table visible with Export button [12] → Click to export data" |
| Action Execution | Execute GUI clicks or API calls via MCP commands | click_input(control_id=12) or execute_word_command("export_table") |
| Result Reporting | Write execution results to shared Blackboard | Write extracted data to subtask_result_1 for HostAgent |
Upon receiving a subtask and execution context from the HostAgent, the AppAgent initializes a ReAct-style control loop where it iteratively:
- Observes the current application state (screenshot + control detection)
- Thinks about the next step (LLM reasoning)
- Acts by executing either a GUI or API-based action (MCP commands)
sequenceDiagram
participant HostAgent
participant AppAgent
participant Application
participant Blackboard
HostAgent->>AppAgent: Delegate subtask<br/>"Extract table from Word"
loop ReAct Loop
AppAgent->>Application: Observe (screenshot + controls)
Application-->>AppAgent: UI state
AppAgent->>AppAgent: Think (LLM reasoning)
AppAgent->>Application: Act (click/API call)
Application-->>AppAgent: Action result
end
AppAgent->>Blackboard: Write result
AppAgent->>HostAgent: Return control
The MCP command system enables reliable control over dynamic and complex UIs by favoring structured API commands whenever available, while retaining fallback to GUI-based interaction commands when necessary.
AppAgent uses a finite state machine with 7 states to control its execution flow:
- CONTINUE: Continue processing the current subtask
- FINISH: Successfully complete the subtask
- ERROR: Encounter an unrecoverable error
- FAIL: Fail to complete the subtask
- PENDING: Wait for user input or clarification
- CONFIRM: Request user confirmation for sensitive actions
- SCREENSHOT: Capture and re-annotate the application screenshot
State Details: See State Machine Documentation for complete state definitions and transitions.
Each execution round follows a 4-phase pipeline:
graph LR
DC[Phase 1:<br/>DATA_COLLECTION<br/>Screenshot + Controls] --> LLM[Phase 2:<br/>LLM_INTERACTION<br/>Reasoning]
LLM --> AE[Phase 3:<br/>ACTION_EXECUTION<br/>GUI/API Action]
AE --> MU[Phase 4:<br/>MEMORY_UPDATE<br/>Record Action]
style DC fill:#e1f5ff
style LLM fill:#fff4e6
style AE fill:#e8f5e9
style MU fill:#fce4ec
Strategy Details: See Processing Strategy Documentation for complete pipeline implementation.
AppAgent executes actions through the MCP (Model-Context Protocol) command system, which provides a unified interface for both GUI automation and native API calls:
# GUI-based command (fallback)
command = Command(
tool_name="click_input",
parameters={"control_id": "12", "button": "left"}
)
await command_dispatcher.execute_commands([command])
# API-based command (preferred when available)
command = Command(
tool_name="word_export_table",
parameters={"format": "csv", "path": "output.csv"}
)
await command_dispatcher.execute_commands([command])Implementation: See Hybrid Actions for details on the MCP command system.
AppAgent is enhanced with Retrieval Augmented Generation (RAG) from heterogeneous sources:
| Knowledge Source | Purpose | Configuration |
|---|---|---|
| Help Documents | Application-specific documentation | Learning from Help Documents |
| Bing Search | Latest information and updates | Learning from Bing Search |
| Self-Demonstrations | Successful action trajectories | Experience Learning |
| Human Demonstrations | Expert-provided workflows | Learning from Demonstrations |
Knowledge Substrate Overview: See Knowledge Substrate for the complete RAG architecture.
AppAgent executes actions through the MCP (Model-Context Protocol) command system:
Application-Level Commands:
capture_window_screenshot- Capture application windowget_control_info- Detect UI controls via UIA/OmniParserclick_input- Click on UI controlset_edit_text- Type text into input fieldannotation- Annotate screenshot with control labels
Command Details: See Command System Documentation for complete command reference.
AppAgent supports multiple control detection backends for comprehensive UI understanding:
UIA (UI Automation):
Native Windows UI Automation API for standard controls
- ✅ Fast and accurate
- ✅ Works with most Windows applications
- ❌ May miss custom controls
OmniParser (Visual Detection):
Vision-based grounding model for visual elements
- ✅ Detects icons, images, custom controls
- ✅ Works with web content
- ❌ Requires external service
Hybrid (UIA + OmniParser):
Best of both worlds - maximum coverage
- ✅ Native controls + visual elements
- ✅ Comprehensive UI understanding
Control Detection Details: See Control Detection Overview.
| Input | Description | Source |
|---|---|---|
| User Request | Original user request in natural language | HostAgent |
| Sub-Task | Specific subtask to execute | HostAgent delegation |
| Application Context | Target app name, window info | HostAgent |
| Control Information | Detected UI controls with labels | Data collection phase |
| Screenshots | Clean, annotated, previous step images | Data collection phase |
| Blackboard | Shared memory for inter-agent communication | Global context |
| Retrieved Knowledge | Help docs, demos, search results | RAG system |
| Output | Description | Consumer |
|---|---|---|
| Observation | Current UI state description | LLM context |
| Thought | Reasoning about next action | Execution log |
| ControlLabel | Selected control to interact with | Action executor |
| Function | MCP command to execute (click_input, set_edit_text, etc.) | Command dispatcher |
| Args | Command parameters | Command dispatcher |
| Status | Agent state (CONTINUE, FINISH, etc.) | State machine |
| Blackboard Update | Execution results | HostAgent |
Example Output:
{
"Observation": "Word document with table, Export button at [12]",
"Thought": "Click Export to extract table data",
"ControlLabel": "12",
"Function": "click_input",
"Args": {"button": "left"},
"Status": "CONTINUE"
}Detailed Documentation:
- State Machine: Complete FSM with state definitions and transitions
- Processing Strategy: 4-phase pipeline implementation details
- Command System: Application-level MCP commands reference
Core Features:
- Hybrid Actions: MCP command system for GUI–API execution
- Control Detection: UIA and visual detection
- Knowledge Substrate: RAG system overview
Tutorials:
- Creating AppAgent: Step-by-step guide
- Help Document Provision: Add help docs
- Demonstration Provision: Add demos
- Wrapping App-Native API: Integrate APIs
:::agents.agent.app_agent.AppAgent
AppAgent Key Characteristics:
✅ Application-Specialized Worker: Dedicated to single Windows application
✅ ReAct Control Loop: Iterative observe → think → act execution
✅ Hybrid Execution: GUI automation + API calls via MCP commands
✅ 7-State FSM: Robust state management for execution control
✅ 4-Phase Pipeline: Structured data collection → reasoning → action → memory
✅ Knowledge-Enhanced: RAG from docs, demos, and search
✅ Orchestrated by HostAgent: Child agent in hierarchical architecture
Next Steps:
- Deep Dive: Read State Machine and Processing Strategy for implementation details
- Learn Features: Explore Core Features for advanced capabilities
- Hands-On Tutorial: Follow Creating AppAgent guide