gemini-3.7-flash is the model Google recommends for it, through the Interactions API.
gemini-2.5-computer-use-preview-10-2025 is now marked legacy. It used the older generate_content call and the types.ComputerUse tool object. The code below uses the current Interactions API.Coordinate space
Orgo computers boot at1280x720x24.
Gemini does not work in pixels. It returns coordinates on a normalized 0-999 grid, scaled to the screenshot you sent. You denormalize every coordinate against the computer’s real resolution before you can click with it:
GET /computers/{id}/screens reports the current width and height.
Setup
Install the required packages. Gemini 3.x computer use needsgoogle-genai 2.7.0 or later:
pip
.env file:
.env
terminal
Complete example
An Orgo computer is a full desktop, so declareenvironment: "desktop". The browser environment adds page-level actions such as navigate and go_back, and expects the current page URL back in every result.
example.py
Usage examples
Basic tasks
Complex workflows
Key concepts
The agent loop
- Request Send the task and a screenshot to the model.
- Actions The model returns
function_callsteps. - Execute Your code runs them in order.
- Screenshot Capture the result and return one
function_resultper call. - Repeat Chain turns with
previous_interaction_iduntil no call comes back.
Reading the response
Everything the model produced for a turn is ininteraction.steps. A step of type function_call carries name, call_id, and arguments. A step of type model_output carries content blocks, and its text blocks are the agent’s final answer.
Every action’s arguments includes an intent string explaining why the model chose it. It is the cheapest debugging signal in the loop, so log it.
System prompt
Pass the Ubuntu guidance assystem_instruction. Without it Gemini single-clicks desktop icons and nothing opens:
Getting your computer ID
Get yourcomputer_id from the Orgo dashboard:
- Go to https://orgo.ai/workspaces
- Click on your project
- Find your computer ID in the computer list
- Use it in:
Computer(computer_id="your-computer-id")
Image format conversion
Orgo returns screenshots as JPEG. Gemini requires PNG, so convert before sending:Safety confirmations
Withenable_prompt_injection_detection on, the model can return a safety_decision of require_confirmation on an action. Prompt your user before running it, and include a safety_acknowledgement in that call’s result when they approve. Skip the confirmation and the action does not run.
Action arguments
Coordinates are normalized 0-999. Denormalizex, y, start_x, start_y, end_x, and end_y before use.
The
desktop environment also exposes mouse_down, mouse_up, key_down, and key_up for held input.
Tool compatibility
drag_and_drop has no SDK helper. Call POST /computers/{id}/drag directly if you need it.
Best practices
1. Clear instructions
2. Send a system instruction, not a user turn
3. Convert coordinates
4. Handle image format
5. Return one result per call
Everyfunction_call needs a matching function_result with the same call_id. Attach the new screenshot to each of them.
6. Add delays
Comparison with Claude and OpenAI
Limitations
- Image format: PNG only, and Orgo returns JPEG, so you convert on every turn.
- Coordinates: normalized, so a wrong
SCREEN_WIDTHorSCREEN_HEIGHTsilently misplaces every click. - Rate limits: subject to Gemini API rate limits.
Troubleshooting
The model does not double-click desktop icons
Pass the Ubuntu system prompt assystem_instruction. Without it, single clicks only select icons.
INVALID_ARGUMENT: Unable to process input image
Gemini received a JPEG. Run the screenshot throughget_screenshot_png() first.
Clicks land in the wrong place
SCREEN_WIDTH and SCREEN_HEIGHT do not match the computer. Check them against GET /computers/{id}/screens.
Missing API key
Ensure both environment variables are set in your.env file:
Next steps
Gemini Docs
Official Gemini computer use documentation
Orgo Quickstart
Learn more about Orgo computers
API Reference
Complete Orgo API documentation