> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.scrapybara.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.scrapybara.com/_mcp/server.

# Act SDK

## What is the Act SDK?

The Act SDK is a unified SDK for building computer use agents with Python and TypeScript. It provides a simple interface for executing looping agentic actions with support for many models and tools. Build production-ready computer use agents with pre-built tools to connect to Scrapybara instances.

## How it works

`act` initiates an interaction loop that continues until the agent achieves its objective. Each iteration of the loop is called a `step`, which consists of the agent's text response, the agent's tool calls, and the results of those tool calls. The loop terminates when the agent returns a message without invoking any tools, and returns `messages`, `steps`, `text`, `output` (if `schema` is provided), and `usage` after the agent's execution.

#### Python

```python
response = client.act(
    model=OpenAI(),
    tools=[
        BashTool(instance),
        ComputerTool(instance),
        EditTool(instance),
    ],
    system=UBUNTU_SYSTEM_PROMPT,
    prompt="Go to the top link on Hacker News",
    on_step=lambda step: print(step.text),
)
messages = response.messages
steps = response.steps
text = response.text
usage = response.usage
```

#### TypeScript

```typescript
const { messages, steps, text, usage } = await client.act({
  model: openai(),
  tools: [
    bashTool(instance),
    computerTool(instance),
    editTool(instance),
  ],
  system: UBUNTU_SYSTEM_PROMPT,
  prompt: "Go to the top link on Hacker News",
  onStep: (step) => console.log(step.text),
});
```

An `act` call consists of 3 core components:

### Model

The model specifies the base LLM for the agent. At each step, the model examines the previous messages, the current state of the computer, and uses tools to take action. Each step will cost an amount of agent credits depending on the model. You can also bring your own API key to bill model charges directly.

#### Python

```python
from scrapybara.openai import OpenAI

model = OpenAI()

# Use your own API key
model = OpenAI(api_key="your_api_key")
```

#### TypeScript

```typescript
import { openai } from "scrapybara/openai";

const model = openai();

// Use your own API key
const model = openai({ apiKey: "your_api_key" });
```

### Tools

Tools are functions that enable agents to interact with the computer. Each tool is defined by a `name`, `description`, and how it can be executed with `parameters` and an execution function. A tool can take in a Scrapybara instance to interact with it directly. Learn more about pre-built tools and how to define custom tools [here](/tools).

#### Python

```python
from scrapybara import Scrapybara
from scrapybara.tools import BashTool, ComputerTool, EditTool

client = Scrapybara()
instance = client.start_ubuntu()

tools = [
    BashTool(instance),
    ComputerTool(instance),
    EditTool(instance),
]
```

#### TypeScript

```typescript
import { ScrapybaraClient } from "scrapybara";
import { bashTool, computerTool, editTool } from "scrapybara/tools";

const client = new ScrapybaraClient();
const instance = await client.startUbuntu();

const tools = [
  bashTool(instance),
  computerTool(instance),
  editTool(instance),
];
```

### Prompt

The prompt is split into two parts, the `system` prompt and a user `prompt`. `system` defines the general behavior of the agent, such as its capabilities and constraints. You can use our provided `UBUNTU_SYSTEM_PROMPT`, `BROWSER_SYSTEM_PROMPT`, and `WINDOWS_SYSTEM_PROMPT` to get started, or define your own. `prompt` should denote the agent's current objective. Alternatively, you can provide `messages` instead of `prompt` to start the agent with a history of messages. `act` conveniently returns `messages` after the agent's execution, so you can reuse it in another `act` call.

#### Python

```python
from scrapybara.prompts import UBUNTU_SYSTEM_PROMPT

system = UBUNTU_SYSTEM_PROMPT
prompt = "Go to the top link on Hacker News"
```

#### TypeScript

```typescript
import { UBUNTU_SYSTEM_PROMPT } from "scrapybara/prompts";

const system = UBUNTU_SYSTEM_PROMPT;
const prompt = "Go to the top link on Hacker News";
```

## Structured output

Use the `schema` parameter to define a desired structured output. The response's `output` field will contain the typed data returned by the model. This is particularly useful when scraping or collecting structured data from websites.

Under the hood, we pass in a `StructuredOutputTool` to enforce and parse the schema.

#### Python

```python
from pydantic import BaseModel
from typing import List

class HNSchema(BaseModel):
    class Post(BaseModel):
        title: str
        url: str 
        points: int
    
    posts: List[Post]

response = client.act(
    model=OpenAI(),
    tools=[
        ComputerTool(instance),
    ],
    schema=HNSchema,
    system=UBUNTU_SYSTEM_PROMPT,
    prompt="Get the top 10 posts on Hacker News",
)

posts = response.output.posts
```

#### TypeScript

```typescript
import { z } from "zod";

const { output } = await client.act({
  model: openai(),
  tools: [
    computerTool(instance),
  ],
  schema: z.object({
    posts: z.array(
      z.object({
        title: z.string(),
        url: z.string(),
        points: z.number(),
      })
    ),
  }),
  system: UBUNTU_SYSTEM_PROMPT,
  prompt: "Get the top 10 posts on Hacker News",
});

const posts = output?.posts;
```

## Agent credits

Consume agent credits or bring your own API key. Without an API key, each step consumes 1 [agent credit](https://scrapybara.com/#pricing). With your own API key, model charges are billed directly to your provider API key.

## Full example

Here is how you can build a computer use agent that can output structured data.

#### Python

```python
from scrapybara import Scrapybara
from scrapybara.openai import OpenAI
from scrapybara.prompts import UBUNTU_SYSTEM_PROMPT
from scrapybara.tools import ComputerTool
from pydantic import BaseModel
from typing import List

client = Scrapybara()
instance = client.start_ubuntu()

class HNSchema(BaseModel):
    class Post(BaseModel):
        title: str
        url: str 
        points: int
    
    posts: List[Post]

response = client.act(
    model=OpenAI(),
    tools=[
        ComputerTool(instance),
    ],
    system=UBUNTU_SYSTEM_PROMPT,
    prompt="Get the top 10 posts on Hacker News",
    schema=HNSchema,
    on_step=lambda step: print(step.text),
)

posts = response.output.posts
print(posts)

instance.stop()
```

#### TypeScript

```typescript
import { ScrapybaraClient } from "scrapybara";
import { openai } from "scrapybara/openai";
import { UBUNTU_SYSTEM_PROMPT } from "scrapybara/prompts";
import { computerTool } from "scrapybara/tools";
import { z } from "zod";

const client = new ScrapybaraClient();
const instance = await client.startUbuntu();

const { output } = await client.act({
  model: openai(),
  tools: [
    computerTool(instance),
  ],
  system: UBUNTU_SYSTEM_PROMPT,
  prompt: "Get the top 10 posts on Hacker News",
  schema: z.object({
    posts: z.array(
      z.object({
        title: z.string(),
        url: z.string(),
        points: z.number(),
      })
    ),
  }),
  onStep: (step) => console.log(step.text),
});

const posts = output?.posts;
console.log(posts);

await instance.browser.stop();
await instance.stop();
```