Skip to main content
Wire Google’s computer use tool to an Orgo computer. Gemini sees a screenshot, returns a batch of actions, and your loop runs them on the computer. Computer use is a built-in tool on Gemini 3.x rather than a separate model. gemini-3.7-flash is the model Google recommends for it, through the Interactions API.
gemini-2.5-computer-use-preview-10-2025 is now marked legacy. It used the older generate_content call and the types.ComputerUse tool object. The code below uses the current Interactions API.

Coordinate space

Orgo computers boot at 1280x720x24. Gemini does not work in pixels. It returns coordinates on a normalized 0-999 grid, scaled to the screenshot you sent. You denormalize every coordinate against the computer’s real resolution before you can click with it:
These two constants are load-bearing. They are the only thing translating Gemini’s grid into pixels. Set them to 1024x768 against a 1280x720 computer and every horizontal click lands roughly 25% short of where Gemini aimed. Nothing errors. The agent looks like it is confidently clicking the wrong thing.
There is no display size to declare on the request. If you create a computer at a non-default resolution, or resize its screen later, set the constants to that resolution instead. GET /computers/{id}/screens reports the current width and height.

Setup

Install the required packages. Gemini 3.x computer use needs google-genai 2.7.0 or later:
pip
Set up your API keys in a .env file:
.env
Or export them as environment variables:
terminal

Complete example

An Orgo computer is a full desktop, so declare environment: "desktop". The browser environment adds page-level actions such as navigate and go_back, and expects the current page URL back in every result.
example.py

Usage examples

Basic tasks

Complex workflows

Key concepts

The agent loop

  1. Request Send the task and a screenshot to the model.
  2. Actions The model returns function_call steps.
  3. Execute Your code runs them in order.
  4. Screenshot Capture the result and return one function_result per call.
  5. Repeat Chain turns with previous_interaction_id until no call comes back.

Reading the response

Everything the model produced for a turn is in interaction.steps. A step of type function_call carries name, call_id, and arguments. A step of type model_output carries content blocks, and its text blocks are the agent’s final answer. Every action’s arguments includes an intent string explaining why the model chose it. It is the cheapest debugging signal in the loop, so log it.

System prompt

Pass the Ubuntu guidance as system_instruction. Without it Gemini single-clicks desktop icons and nothing opens:

Getting your computer ID

Get your computer_id from the Orgo dashboard:
  1. Go to https://orgo.ai/workspaces
  2. Click on your project
  3. Find your computer ID in the computer list
  4. Use it in: Computer(computer_id="your-computer-id")

Image format conversion

Orgo returns screenshots as JPEG. Gemini requires PNG, so convert before sending:

Safety confirmations

With enable_prompt_injection_detection on, the model can return a safety_decision of require_confirmation on an action. Prompt your user before running it, and include a safety_acknowledgement in that call’s result when they approve. Skip the confirmation and the action does not run.

Action arguments

Coordinates are normalized 0-999. Denormalize x, y, start_x, start_y, end_x, and end_y before use. The desktop environment also exposes mouse_down, mouse_up, key_down, and key_up for held input.

Tool compatibility

drag_and_drop has no SDK helper. Call POST /computers/{id}/drag directly if you need it.

Best practices

1. Clear instructions

2. Send a system instruction, not a user turn

3. Convert coordinates

4. Handle image format

5. Return one result per call

Every function_call needs a matching function_result with the same call_id. Attach the new screenshot to each of them.

6. Add delays

Comparison with Claude and OpenAI

Limitations

  • Image format: PNG only, and Orgo returns JPEG, so you convert on every turn.
  • Coordinates: normalized, so a wrong SCREEN_WIDTH or SCREEN_HEIGHT silently misplaces every click.
  • Rate limits: subject to Gemini API rate limits.

Troubleshooting

The model does not double-click desktop icons

Pass the Ubuntu system prompt as system_instruction. Without it, single clicks only select icons.

INVALID_ARGUMENT: Unable to process input image

Gemini received a JPEG. Run the screenshot through get_screenshot_png() first.

Clicks land in the wrong place

SCREEN_WIDTH and SCREEN_HEIGHT do not match the computer. Check them against GET /computers/{id}/screens.

Missing API key

Ensure both environment variables are set in your .env file:

Next steps

Gemini Docs

Official Gemini computer use documentation

Orgo Quickstart

Learn more about Orgo computers

API Reference

Complete Orgo API documentation