# iPhone Use: How to Let OpenAI Codex Control a Real iPhone with AI

Learn how iPhone Use lets Codex control a real iPhone via USB, MCP and WebDriverAgent. Explore installation, 17 tools, features, security and alternatives.

Canonical URL: https://aiidelist.com/blog/iphone-use-codex-control-real-iphone

Language: en

Published: 2026-10-09

Updated: 2026-10-09

## Key Takeaways

- **iPhone Use** is an open-source Codex plugin that enables AI agents to interact with a real iPhone through natural-language instructions. It supports opening apps, reading screens, tapping, swiping, entering text, and collecting information.
- Developed by **zhongerxin (Twox / 钟二信)**, the project was created on October 6, 2026, and had approximately 1,300 GitHub stars as of October 9, 2026.
- The project uses **WebDriverAgent (WDA), USB communication, and the Model Context Protocol (MCP)** to connect OpenAI Codex with native iOS interfaces.
- It provides **17 AI-accessible tools and two additional screen-widget tools**, along with setup guidance and a live iPhone screen preview.
- The latest verified release, **v0.3.7, published October 9, 2026**, improves accessibility-tree processing and introduces documented anonymous usage analytics.
- iPhone Use is free under the MIT license. However, real-device automation requires macOS, Xcode, a compatible iPhone, and a properly signed WebDriverAgent Runner.
- Its strongest applications include mobile app testing, repetitive iPhone workflows, UI-based data collection, and AI-assisted interaction with apps that do not expose suitable APIs.

![Image](https://cdn.aiidelist.com/api/image/s011iCkHCFqddMBlDpsri.webp)

## What Is iPhone Use?

**iPhone Use is an open-source AI automation tool that allows OpenAI Codex to operate a physical iPhone over USB.** Instead of writing traditional automation scripts for every interaction, developers can describe their goals in natural language and let an AI agent navigate the phone's interface.

The project is maintained by zhongerxin on GitHub. Its central idea is straightforward: give a coding agent access to the same kind of visual and interactive feedback that a person receives when using an iPhone.

For example, a developer could instruct Codex to open Notes, create a draft, enter a paragraph, and verify that the content appears correctly. Another workflow might involve opening an application, searching for specific information, scrolling through results, and returning a structured summary.

Unlike an API integration, iPhone Use interacts with the actual application interface. This makes it potentially useful for native iOS applications that provide no external automation API.

However, iPhone Use is not a standalone artificial intelligence model or an unrestricted remote-control service. It is an automation layer that allows a compatible AI agent to observe and operate a connected device. The reasoning and task planning remain the responsibility of the AI agent.

![iPhone Use showing a real iPhone screen inside Codex](https://raw.githubusercontent.com/zhongerxin/iPhone-use/main/assets/iphone-use-demo.png)

### iPhone Use at a Glance

| Attribute | Details |
|---|---|
| Project | iPhone Use |
| Developer | zhongerxin (Twox / 钟二信) |
| Repository | zhongerxin/iPhone-use |
| First created | October 6, 2026 |
| Latest verified release | v0.3.7, October 9, 2026 |
| Primary language | Python |
| License | MIT |
| GitHub popularity | Approximately 1.3K stars as of October 9, 2026 |
| Primary AI integration | OpenAI Codex |
| Communication | USB and local MCP server |
| iOS automation engine | WebDriverAgent / XCTest |
| Model-accessible tools | 17 |
| Additional widget tools | 2 |
| Supported target | Physical iPhone |
| Supported host | macOS with full Xcode |
| Pricing | Free, open-source software |

GitHub popularity and release versions can change rapidly. The repository releases provide the latest published version information.

## How Does iPhone Use Work?

The most important technical distinction is that iPhone Use does not magically gain direct access to iOS applications. It relies on Apple's existing UI testing infrastructure, exposed through WebDriverAgent.

The system has four main layers:

1. **Codex:** Interprets natural-language instructions, plans actions, and evaluates results.
2. **MCP server:** Exposes phone automation tools to the AI agent and coordinates requests.
3. **WebDriverAgent:** Executes UI automation commands through Apple's XCTest infrastructure.
4. **Physical iPhone:** Runs the actual applications and returns interface observations.

The typical execution pipeline looks like this:

```text
User instruction
       |
       v
OpenAI Codex
       |
       v
MCP tool invocation
       |
       v
iPhone Use local server
       |
       v
USB forwarding + WebDriverAgent
       |
       v
Physical iPhone
       |
       v
UI observation / screenshot
       |
       v
Codex interprets the result
       |
       v
Next action or final response
```

### Why WebDriverAgent Matters

WebDriverAgent is an established component of the iOS automation ecosystem. It communicates with applications through Apple's XCTest framework and makes interactions such as tapping, typing, and reading interface elements available to external tools.

For a physical iPhone, the runner must be properly signed and installed on the device. This is necessary because iOS restricts which processes can execute automated UI testing operations.

The result is more sophisticated than simply streaming screenshots and clicking approximate screen coordinates.

When an application exposes accessible UI elements, the automation system can identify controls using structured information such as element type, text, position, and visibility.

When an element cannot be located reliably, iPhone Use can return a screenshot and allow the AI agent to select a coordinate-based interaction instead.

This hybrid approach is particularly important for applications containing custom interfaces, graphical controls, or incomplete accessibility metadata.

### Why MCP Matters

The Model Context Protocol provides a standardized way for AI applications to discover and invoke external capabilities.

With MCP, Codex can call tools such as `pua_observe`, `pua_tap`, and `pua_type_text` without embedding iPhone automation code directly into every conversation.

This separation allows the AI model to focus on planning and interpretation while the local automation service handles device communication, execution, and error reporting.

## What Can iPhone Use Do? Key Features Explained

### 1. Control Native iPhone Applications

Codex can launch installed applications, navigate visible interfaces, interact with controls, and move between screens.

Potential tasks include opening a productivity app, finding a particular item, creating a draft, changing an accessible setting, or inspecting an application's navigation flow.

This is especially interesting for developers because many native apps cannot be meaningfully tested through browser automation alone.

### 2. Read iPhone Screens and UI Elements

The `pua_observe` tool can retrieve interface information, including accessibility elements and screenshots.

Structured observations are generally more useful than screenshots alone because they expose information that the model can use to locate and distinguish controls.

For instance, a screen may contain several visually similar buttons. An accessibility snapshot can provide labels, types, and coordinates that help the agent choose the correct target.

However, not every iOS interface exposes a complete accessibility tree. Custom-rendered surfaces and protected screens may require visual inspection or may remain inaccessible.

### 3. Tap, Swipe, Scroll, and Use Device Buttons

The plugin supports common interaction patterns through tools including `pua_tap`, `pua_swipe`, and `pua_press_button`.

These actions allow an agent to navigate lists, open menus, interact with visible controls, and return to the Home screen.

The distinction between actions and verified results matters. A successful command response does not necessarily establish that the application reached the intended state.

For critical actions, the agent should observe the resulting interface before declaring completion.

### 4. Enter Chinese, Unicode, and Long Text

The `pua_type_text` tool supports Unicode input and longer text-entry operations.

This is useful for note-taking, search fields, multilingual testing, and applications that require substantial text input.

Long-text handling includes chunking and continuation mechanisms, reducing the risk of losing an entire input operation when interruptions occur.

The plugin separates text entry from submission by default. Entering text into a field does not automatically mean pressing Send, Save, or Confirm.

### 5. Search and Collect Long Lists

Many mobile interfaces load additional information only as users scroll.

The `pua_scroll_find` and `pua_collect_list` tools help an agent work through these interfaces using bounded scrolling, overlapping observations, and deduplication.

For example, an agent may search a long settings list for a specific option or collect visible entries from a paginated application screen.

This does not provide unrestricted access to application databases. Collection remains limited to information exposed through the interface and the user's authorized access.

### 6. Batch Multiple Actions

The `pua_batch` tool can combine known actions into a single tool invocation.

Instead of asking the AI model to issue separate calls for every predictable step, a batch can execute several actions under a shared execution budget.

This reduces model round trips and coordination overhead.

Batching works best when the next actions are already known. Unknown screens, confirmation dialogs, and unexpected navigation changes should trigger a new observation rather than blind continuation.

### 7. Live iPhone Screen Preview

One of the project's most distinctive features is its live screen widget inside Codex.

The interface displays the connected iPhone within a simulated device frame. It can show connection status, visible interactions, and animated indicators for taps or drags.

The widget also provides controls for refreshing the connection, returning to the Home screen, and capturing a screenshot.

The preview uses a separate screen-streaming mechanism rather than requiring a full accessibility snapshot for every displayed frame.

This makes it easier for users to understand what the agent is doing while maintaining a separate observation channel for the model.

### 8. Connection Recovery and Failure Handling

Mobile automation frequently encounters transient errors: an app may take longer to load, a previous element reference may become invalid, or the device may disconnect.

iPhone Use attempts to make these failures explicit rather than hiding them behind repeated operations.

When an action has an uncertain outcome, the agent is expected to inspect the current state before attempting another mutation.

This matters for operations involving text submission, data modification, purchases, and other actions that should not be performed twice accidentally.

## The 17 iPhone Use Tools

The project currently exposes 17 model-accessible tools, plus two additional tools reserved for its screen widget.

| Tool | Main purpose |
|---|---|
| `pua_doctor` | Diagnose local dependencies and connection problems |
| `pua_setup` | Configure the iPhone environment and manage setup jobs |
| `pua_ready` | Verify that the device automation environment is ready |
| `pua_observe` | Read interface information and screenshots |
| `pua_find` | Find controls or matching interface elements |
| `pua_apps` | Query application information and identifiers |
| `pua_launch_app` | Launch an application |
| `pua_tap` | Tap an element or screen coordinate |
| `pua_swipe` | Execute swipe and drag interactions |
| `pua_press_button` | Perform supported device-button operations |
| `pua_type_text` | Enter Unicode text |
| `pua_wait` | Wait for specified interface conditions |
| `pua_batch` | Combine several actions |
| `pua_scroll_find` | Search for an element while scrolling |
| `pua_collect_list` | Collect and deduplicate visible list entries |
| `pua_screen` | Manage the screen preview |
| `pua_metrics` | Inspect bounded operation timing and usage measurements |

The `pua_` naming convention was standardized in version 0.3.6. Earlier versions used `wda_` prefixes, so older installation instructions and demonstrations may contain outdated tool names.

The complete tool list is documented in the project's official README.

## How to Install iPhone Use: Step-by-Step Tutorial

Installing iPhone Use requires more preparation than installing a conventional browser extension. The principal challenge is configuring Apple's real-device development and code-signing environment.

### Step 1: Check the System Requirements

Before installation, confirm the following:

| Requirement | Details |
|---|---|
| Mac | macOS computer |
| Xcode | Full Xcode installation, with first-launch setup completed |
| iPhone | Physical iPhone connected through USB |
| Apple account | Available in Xcode with a usable development team |
| Developer Mode | Enabled when required by the device |
| Python | 3.9 or newer |
| Node.js | 20.19+, 22.12+, or 24+ supported branches |
| npm | 10 or newer |
| Codex | Codex desktop application or CLI |

Xcode must support the iOS version installed on the device. Installing Command Line Tools alone is insufficient for the required WDA build and signing process.

A jailbreak is not required, and the project does not require a separate Appium Server process.

### Step 2: Clone the GitHub Repository

Open Terminal on the Mac and run:

```bash
git clone https://github.com/zhongerxin/iPhone-use.git
cd iPhone-use
```

Before executing installation scripts from any third-party repository, review their behavior and permissions.

### Step 3: Install the Codex Plugin

From the repository directory, run:

```bash
sh scripts/install.sh
```

According to the project documentation, the installer validates and stages the source, registers the local Codex plugin marketplace, installs the plugin and its skills, and configures the MCP server under the `iphone_use` namespace.

After installation, restart or reconnect the Codex conversation so the newly registered tools become available.

### Step 4: Connect and Prepare the iPhone

Connect the iPhone using a USB data cable.

Unlock the device and accept the computer trust prompt. When necessary, enable Developer Mode from the iPhone's Privacy & Security settings and complete the required restart.

Open Xcode and confirm that the device appears as an available development target.

The WDA Runner must be signed using a development team and bundle identifier that are valid for the user's own environment.

Do not copy another developer's device identifiers, signing information, or provisioning configuration.

### Step 5: Let Codex Configure WebDriverAgent

The project includes an `iphone-use-setup` skill designed to guide the initial configuration.

In a new Codex conversation, enter:

```text
Use iphone-use-setup to configure my USB-connected iPhone.
Check the required dependencies and device connection.
Reuse an existing healthy WebDriverAgent installation if available.
Otherwise, configure signing with my own Apple development team,
build and start WebDriverAgent Runner, and verify readiness.
When pua_ready returns ready=true, show the live phone screen.
```

The setup workflow uses a pinned WebDriverAgent 16.14.0 revision and manages the download, configuration, build, and startup process.

Some actions still require direct user participation, including Apple account authentication, trust confirmation, enabling Developer Mode, and unlocking the device.

### Step 6: Verify the Connection

The important readiness condition is:

```text
pua_ready -> ready=true
```

A successful build is not enough. The WDA service must also be running and accessible, and the plugin must establish a working device session.

Only after the readiness check succeeds should the agent begin normal phone operations.

### Step 7: Run the First Automation Task

Start with a reversible task that does not require sensitive information.

For example:

```text
Open the Notes app on my connected iPhone.
Create a new note containing the text:
Testing iPhone Use with Codex.
Verify that the text appears correctly.
Do not delete existing notes or share the new note.
```

This simple test checks several capabilities at once: application launch, screen navigation, text input, and result verification.

## Practical iPhone Use Examples

### Example 1: Test an iOS Application

For iOS developers, one of the most practical applications is testing a user journey on a real device.

```text
Open the application currently under development.
Navigate to its settings page.
Check whether the expected controls are visible.
Open the profile editor and enter sample text.
Verify the updated interface.
Report any unexpected screen or navigation behavior.
Do not modify production account settings.
```

Compared with a static screenshot review, this workflow provides feedback from the actual running application.

### Example 2: Collect Information from a List

```text
Open the selected application and navigate to its item list.
Read the visible item names.
Scroll through the list with overlapping observations.
Deduplicate entries.
Return the collected information as a structured table.
Explain which part of the list was covered and whether any entries
could not be verified.
```

This example demonstrates why bounded collection and explicit coverage reporting are important. An agent should not claim to have collected an entire list if it only observed several pages.

### Example 3: Draft Content in a Mobile App

```text
Open the Notes app.
Create a new draft containing a short project summary.
Check the title and body for accuracy.
Leave the draft available for manual review.
Do not send, publish, or share anything.
```

The separation between drafting and submitting reduces the risk of unintended external actions.

## What Changed in iPhone Use v0.3.7?

The October 9, 2026 release focused on improving accessibility processing and documenting usage analytics.

One important optimization was removing an unused `accessible` attribute from ordinary full-page XML observations.

According to the developer's v0.3.7 release notes, real-device comparisons involving WeChat, Taobao, and the Photos interface showed equivalent UI-node information before and after the change.

The raw XML size decreased by approximately **7%–8%** in those comparisons.

This improvement is relevant because structured UI snapshots can become large and expensive to process, particularly when pages contain many repeated elements.

However, a smaller XML response does not guarantee a proportional reduction in total task latency.

The release notes explicitly acknowledge that snapshot performance still depends on the interface and caching conditions. Full-page observations can continue to time out, including in complex message-list interfaces.

The release also includes a documented PostHog integration for anonymous usage and reliability metrics, with an opt-out option.

These changes reflect an emphasis on reducing unnecessary automation overhead while retaining enough information for the AI agent to make reliable decisions.

## Performance: Why Structured UI Observation Beats Blind Clicking

There are two common approaches to AI-driven device interaction.

**Screenshot-first automation** repeatedly captures the screen, asks a vision model to identify a target, and executes a coordinate-based action.

**Accessibility-first automation** reads structured UI elements and uses them to select controls. Screenshots become a fallback when the structured information is incomplete.

iPhone Use primarily follows the second approach.

This has several practical benefits:

- Structured observations can reduce the amount of visual data that needs to be interpreted.
- Element attributes help distinguish similarly positioned controls.
- Explicit identifiers and text labels can make target selection more stable.
- Combining known actions reduces repeated model-tool communication.
- Screenshot fallback remains available for custom or poorly labeled interfaces.

The trade-off is that accessibility snapshots themselves can be slow or incomplete. Large interfaces, dynamic layouts, and XCTest limitations can still cause delays.

The project also uses connection reuse, compact output, operation locks, and bounded waits to reduce unnecessary work.

For performance-sensitive applications, useful measurements include time to readiness, screenshot latency, accessibility snapshot latency, task completion time, and the rate of failed or repeated actions.

No independently verified universal task-success benchmark has been established for iPhone Use, so performance claims should be evaluated on representative applications rather than assumed from individual demonstrations.

## Common iPhone Use Problems and How to Fix Them

### iPhone Not Detected

If the device does not appear in the automation environment, first check the USB cable, device trust, and whether the phone is unlocked.

Confirm that Xcode recognizes the device before troubleshooting the AI plugin itself.

### Developer Mode Is Disabled

Recent iOS versions require Developer Mode for development-related operations.

Enable it on the device, complete the required restart, and repeat device discovery.

Depending on the configuration, Apple's UI Automation setting may also need to be enabled. The Appium real-device preparation guide provides additional details.

### WDA Signing Fails

Code signing is one of the most common setup obstacles.

Check the selected Xcode development team, bundle identifier, signing certificate, and provisioning profile.

A free Apple development team can support manual WDA signing in appropriate configurations, although free provisioning profiles have additional limitations and may require periodic renewal.

A paid Apple Developer membership is not universally required simply to try the tool.

### Build Succeeds but the Phone Is Not Ready

A successful WDA build does not prove that the automation service is responding.

Inspect the setup status and check that the Runner remains active, USB forwarding works, and `pua_ready` returns a positive readiness result.

Avoid repeatedly rebuilding an already valid installation when the actual problem is device communication.

### Controls Cannot Be Found

A custom-rendered interface may expose limited accessibility information.

Use a fresh screenshot and inspect the visible target. If appropriate, switch to coordinate-based tapping.

The project documents a `image.pixel_to_point` conversion factor for translating screenshot pixels into iPhone point coordinates. Ignoring this difference can produce incorrect taps on high-density displays.

### An Action Times Out After Submission

This is more serious than a simple connection failure.

The action may have completed on the device even though the agent did not receive a successful response.

For actions involving messages, purchases, data entry, or saving changes, inspect the actual application state before retrying.

Repeating the entire operation without verification can create duplicate content or unintended transactions.

### iPhone Mirroring Conflicts

The project's troubleshooting documentation identifies possible conflicts with macOS iPhone Mirroring when the automation system encounters unavailable or incomplete interface data.

If a conflict is confirmed, exit iPhone Mirroring, unlock the device, and rerun the readiness check.

The presence of the mirroring process alone is not proof that it caused the error.

More detailed recovery guidance is available in the project's troubleshooting documentation.

## Is iPhone Use Safe? Privacy and Security Considerations

An AI agent with the ability to operate a physical phone deserves more security scrutiny than a conventional code-completion extension.

Depending on the application and the user's permissions, a connected agent may encounter messages, account information, private files, financial interfaces, and other sensitive content.

### Local Communication

The project uses a local USB connection and loopback-based communication with WDA. Its documentation states that runtime configuration and signing data remain on the host Mac.

The default forwarded port is 18100, and the project's remote WDA URL restrictions are designed to prevent arbitrary remote endpoints.

However, local device communication should not be confused with a guarantee that every component involved in the workflow operates offline. The AI model provider, host environment, and optional telemetry have separate data-handling considerations.

### Anonymous Analytics

As of v0.3.7, anonymous PostHog analytics are enabled by default.

According to the project's analytics documentation, telemetry includes bounded information about plugin sessions, tool usage, connection states, durations, and error categories.

The developer states that phone screenshots, input text, accessibility trees, clipboard contents, application bundle identifiers, device identifiers, and signing details are excluded from telemetry.

These statements describe the project's documented behavior and should not be interpreted as an independent security audit.

Users who prefer to disable this analytics collection can configure the MCP server environment:

```toml
[mcp_servers.iphone_use.env]
IPHONE_USE_ANALYTICS = "0"
```

The configuration belongs in the Codex MCP settings, typically `~/.codex/config.toml`. Restart the MCP service afterward.

Alternatively, the environment variable `DO_NOT_TRACK=1` is supported.

### Recommended Security Practices

- Inspect third-party installation scripts before execution.
- Keep device automation interfaces restricted to trusted local processes.
- Handle passwords, verification codes, and biometric authentication directly on the phone.
- Require explicit confirmation before purchases, publishing, deleting data, or sending messages.
- Avoid granting an agent access to sensitive applications unless the task requires it.
- Review screenshots and logs before sharing troubleshooting information.
- Use a dedicated development or test device for experiments involving unfamiliar workflows.

The official workflow allows users to take over during authentication, with preview and phone interactions paused until the user completes the necessary step.

## iPhone Use vs Mobile MCP vs agent-device

Although iPhone Use is a notable entry in AI-driven mobile automation, it is not the only tool in this category.

Two relevant alternatives are Mobile Next's Mobile MCP and Callstack's agent-device.

| Feature | iPhone Use | Mobile MCP | agent-device |
|---|---|---|---|
| Primary focus | Codex-controlled real iPhone | Cross-platform mobile automation | AI-assisted application testing and verification |
| iOS support | Real iPhone | Real devices and simulators | Real devices and simulators |
| Android support | Not a documented target | Yes | Yes |
| MCP integration | Yes | Yes | Yes |
| Structured UI observation | Yes | Yes | Yes, where supported |
| Screenshot-based interaction | Yes | Yes | Yes |
| Device preview | Codex screen widget | Available mirroring integrations | Device and evidence workflows |
| Main advantage | Focused Codex integration and guided setup | Unified iOS/Android automation | Testing, replay, CI, and broader development workflows |

### When to Choose iPhone Use

Choose iPhone Use when the primary objective is connecting a physical iPhone to Codex with an integrated setup workflow and visible screen feedback.

Its narrower platform focus can be an advantage for users who specifically want to explore real-iPhone automation rather than maintain a cross-platform testing infrastructure.

### When to Choose Mobile MCP

Mobile MCP is better suited to projects requiring automation across iOS and Android devices, including simulators, emulators, and physical hardware.

Its platform-independent interface is valuable for teams working with multiple mobile operating systems.

### When to Choose agent-device

agent-device is particularly relevant when the goal is to let AI coding agents verify application changes, capture evidence, save repeatable automation flows, and integrate checks into development pipelines.

For professional mobile development teams, deterministic replay and CI integration may be more valuable than an individual assistant controlling a phone interactively.

These products overlap, but their intended workflows are different. The best choice depends on whether the priority is personal device interaction, cross-platform automation, or software verification.

## Advantages and Limitations of iPhone Use

### Advantages

- **Real-device interaction:** Operations run against an actual iPhone rather than only a simulated interface.
- **Natural-language control:** Developers can describe tasks instead of scripting every UI interaction manually.
- **Structured observation:** Accessibility data helps agents reason about the current application state.
- **Visual fallback:** Screenshots support recovery when structured element targeting fails.
- **Live preview:** Users can observe the device while the agent works.
- **Reduced repetition:** Batch operations and connection reuse can lower orchestration overhead.
- **Open source:** The MIT license allows inspection, modification, and redistribution under its terms.

### Limitations

- **Initial setup complexity:** Xcode, signing, provisioning, and WDA deployment remain important barriers.
- **Mac dependency:** The documented installation workflow requires macOS.
- **Physical iPhone focus:** The project is not presented as a complete Android automation platform.
- **Interface variability:** Custom controls and changing layouts may reduce reliability.
- **Task uncertainty:** Complex multi-step tasks can still fail or require user intervention.
- **Authentication boundaries:** Passwords and biometric checks require appropriate human participation.
- **Early-stage maturity:** Rapid releases provide improvements but can also change tool names, configuration, and runtime behavior.

## Frequently Asked Questions

### Can Codex Really Control an iPhone?

Yes. With iPhone Use configured and WebDriverAgent running on a connected device, Codex can perform supported UI interactions through MCP tools.

The agent still depends on accessible application interfaces, valid device permissions, and successful execution feedback.

### Does iPhone Use Require Jailbreaking?

No. The documented approach uses Apple's XCTest automation infrastructure and a properly signed WebDriverAgent Runner.

### Does iPhone Use Work Without a Mac?

The official installation and deployment workflow requires macOS and full Xcode. Other remote-device architectures may be possible, but they are not the standard setup documented by this project.

### Does iPhone Use Require a Paid Apple Developer Account?

Not necessarily. Free Apple accounts can support manual development signing under Apple's restrictions. The account must still provide a valid development team and provisioning profile for the WDA Runner.

### Is iPhone Use Free?

Yes. The GitHub project is available under the MIT license. Users may still incur costs associated with hardware, development infrastructure, and any AI model service used with Codex.

### Can iPhone Use Read Messages or Interact with Social Apps?

Potentially, if the user has authorized access and the application exposes the necessary information through its interface.

Successful interaction is not guaranteed for every app. Login requirements, dynamic layouts, custom controls, and platform restrictions may interrupt automation.

### Can iPhone Use Perform Purchases Automatically?

The underlying UI automation tools may interact with ordinary buttons and forms, but financial transactions introduce significant reliability and security risks.

Sensitive or irreversible actions should require explicit user authorization and final-state verification. A successful button tap alone is insufficient evidence that a transaction completed correctly.

### Can Claude Code Use iPhone Use?

The project is specifically packaged and documented for Codex. Because it exposes an MCP server, integration with another compatible client may be technically possible, but the official installation and widget experience should not be assumed to work identically in Claude Code.

For broader client support, Mobile MCP and agent-device are relevant alternatives.

### Is iPhone Use the Same as iPhone Mirroring?

No. Apple's iPhone Mirroring provides a human-facing way to interact with an iPhone from a Mac. iPhone Use instead exposes structured automation operations to an AI agent through WebDriverAgent and MCP.

The project's documentation also identifies situations where the two approaches can conflict.

### What Is the Latest iPhone Use Version?

As of October 9, 2026, the latest verified GitHub release is v0.3.7. The version adds accessibility-processing improvements and includes documented anonymous analytics controls. The GitHub releases page should be checked for subsequent updates.

## The Bigger Picture: AI Agents Are Moving Beyond Code Editors

The wider significance of iPhone Use extends beyond controlling a single device.

Coding agents have traditionally been evaluated by how well they generate code, execute terminal commands, edit repositories, and troubleshoot development environments.

Mobile automation adds another capability: interacting with software through the same interfaces used by ordinary users.

This creates new opportunities for AI-assisted quality assurance, repetitive mobile workflows, accessibility-oriented interaction, and application testing without custom integration code for every task.

The central challenge is reliability.

A useful mobile agent must do more than select the correct button. It must recognize the current state, distinguish completed actions from uncertain ones, recover from interruptions, avoid duplicate submissions, and provide evidence that the requested outcome occurred.

iPhone Use addresses several of these challenges through bounded actions, structured observations, screenshot fallback, and explicit recovery semantics.

Its long-term significance will depend less on the novelty of watching an AI tap an iPhone and more on whether developers can turn those interactions into dependable, repeatable workflows.

## Conclusion

**iPhone Use demonstrates how OpenAI Codex can move beyond code generation and interact directly with applications running on a physical iPhone.** By combining WebDriverAgent, local USB communication, MCP tools, and a live screen interface, the project turns existing iOS testing infrastructure into an AI-accessible automation system.

Its 17 model-accessible tools cover app navigation, interface observation, text input, gestures, list collection, batching, and readiness checks. The project also includes recovery mechanisms that help agents distinguish failed actions from uncertain outcomes.

The main obstacles remain initial device setup, code signing, changing application interfaces, and the difficulty of verifying complex tasks reliably.

For Codex users interested in real-device automation, iPhone Use is a promising open-source project to explore. For teams building larger mobile automation systems, it is also worth evaluating alongside Mobile MCP and agent-device.

**To get started, visit the official iPhone Use GitHub repository, review the installation instructions, connect a compatible iPhone, and begin with a small, reversible automation task.**
