Skip to navigation

Batch Operations

The batch operate SDK helpers allow you to efficiently process multiple URLs in parallel while managing browser sessions and windows automatically. These helpers automatically manage the lifecycle of browser sessions and windows, handling creation and cleanup behind the scenes so you can focus on your own application logic.

This guide explains how to use these helpers which are available in both the Node.js and Python Airtop SDKs.

Python SDK

Basic Usage

from airtop import AsyncAirtop, types, BatchOperationUrl, BatchOperationInput, BatchOperationResponse, BatchOperateConfig
from dataclasses import dataclass
from typing import Optional, List, Any, Dict
async def main():
client = AsyncAirtop(api_key="your-api-key")
# Define URLs to process
urls = [
BatchOperationUrl(url="https://example.com/page1"),
BatchOperationUrl(url="https://example.com/page2"),
BatchOperationUrl(url="https://example.com/page3")
]
# Define your operation function
async def operation(input: BatchOperationInput) -> BatchOperationResponse:
# Example: Run a custom query on the page
result = await client.windows.page_query(
session_id=input.session_id,
window_id=input.window_id,
prompt="What is the main idea of this page?"
)
return BatchOperationResponse(
data=result.data.model_response,
should_halt_batch=False,
additional_urls=[]
)
# Execute batch operation
results = await client.batch_operate(
urls=urls,
operation=operation,
)
# Results will be an array containing the responses from each operation
# Example:
# [
# "The main idea of this page is about cloud computing services and their benefits for enterprise...",
# "This page discusses machine learning algorithms and their applications in data science...",
# "The page focuses on cybersecurity best practices for organizations..."
# ]

Advanced Features

Error Handling

By default, the Batch Operate helpers will not automatically retry failed operations. However, you can implement this or other custom error handling by providing an on_error callback.

async def handle_error(error: BatchOperationError):
print(f"Error occurred: {error.error}")
print(f"Session ID: {error.session_id}")
print(f"Window ID: {error.window_id}")
print(f"URLs affected: {error.operation_urls}")
config = BatchOperateConfig(
on_error=handle_error
)

Session Configuration

You can customize the session configuration for all sessions created during the batch operation.

A very helpful option is the profile_name parameter. This allows you to re-use a browser Profile that has sign-in credentials and cookies for websites that require authentication.

session = await client.sessions.create()
# ... Your custom logic, i.e. authentication ...
# Terminate the session to persist the profile
await client.sessions.terminate(session.data.id)
config = BatchOperateConfig(
session_config=types.SessionConfigV1(
timeout_minutes=10, # Terminate the session after 10 mins of inactivity
profile_name="my-linkedin-profile" # Use a profile that has been signed in to LinkedIn
)
)
results = await client.batch_operate(
urls=urls,
operation=operation,
config=config
)

Halting Operations

You can stop the batch processing at any point by returning should_halt_batch=True:

async def operation(input: BatchOperationInput) -> BatchOperationResponse:
# ... your custom logic ...
if some_condition:
return BatchOperationResponse(
data=result,
should_halt_batch=True # This will stop processing remaining URLs
)
return BatchOperationResponse(data=result)

Controlling Concurrency

Note: We recommend keeping max_windows_per_session=1 (the default) for optimal stability and performance. While you can experiment with multiple windows per session, single-window sessions are generally more reliable and easier to manage. Increase max_concurrent_sessions instead if you need more parallelization.

You can control the number of concurrent sessions and windows per session:

config = BatchOperateConfig(
max_concurrent_sessions=30, # Default value is 30
max_windows_per_session=1 # Default value is 1
)

Dynamic URL Addition

Your operation can discover and add new URLs to process during execution. The original batch_operate call will not complete until all operations have completed, including the new additions.

async def operation(input: BatchOperationInput) -> BatchOperationResponse:
# ... your scraping logic ...
new_urls = [
BatchOperationUrl(url="https://example.com/discovered1"),
BatchOperationUrl(url="https://example.com/discovered2")
]
return BatchOperationResponse(
data=result,
additional_urls=new_urls
)

Node.js SDK

Basic Usage

import {
Airtop,
BatchOperationUrl,
BatchOperationInput,
BatchOperationResponse,
BatchOperateConfig
} from '@airtop/sdk';
async function main() {
const client = new Airtop({ apiKey: "your-api-key" });
// Define URLs to process
const urls: BatchOperationUrl[] = [
{ url: "https://example.com/page1" },
{ url: "https://example.com/page2" },
{ url: "https://example.com/page3" }
];
// Define your operation function
const operation = async (input: BatchOperationInput): Promise<BatchOperationResponse> => {
const { windowId, sessionId } = input;
// Example: Run a custom query on the page
const result = await client.windows.pageQuery({
sessionId,
windowId,
prompt: "What is the main idea of this page?"
});
return {
data: result.data.modelResponse,
shouldHaltBatch: false,
additionalUrls: []
};
};
// Execute batch operation
const results = await client.batchOperate({
urls,
operation
});
// Results will be an array containing the responses from each operation
// Example results:
// [
// "This page is about cloud computing services and infrastructure",
// "The main topic is machine learning and AI applications",
// "This page discusses data analytics and visualization tools"
// ]
}

Advanced Features

Error Handling

By default, the Batch Operate helpers will not automatically retry failed operations. However, you can implement this or other custom error handling by providing an onError callback.

const config: BatchOperateConfig = {
onError: async (error: BatchOperationError) => {
console.error(`Error occurred: ${error.error}`);
console.error(`Session ID: ${error.sessionId}`);
console.error(`Window ID: ${error.windowId}`);
console.error(`URLs affected:`, error.operationUrls);
}
};

Session Configuration

You can customize the session configuration for all sessions created during the batch operation.

A very helpful option is the profileName parameter. This allows you to re-use a browser Profile that has sign-in credentials and cookies for websites that require authentication.

// Create and setup initial session
const session = await client.sessions.create();
// ... Your custom logic, i.e. authentication ...
// Terminate the session to persist the profile
await client.sessions.terminate({ sessionId: session.data.id });
const config: BatchOperateConfig = {
sessionConfig: {
timeoutMinutes: 10, // Terminate the session after 10 mins of inactivity
profileName: "my-linkedin-profile", // Use a profile that has been signed in to LinkedIn
}
};
const results = await client.batchOperate({
urls,
operation,
config
});

Halting Operations

You can stop the batch processing at any point by returning shouldHaltBatch: true:

const operation = async (input: BatchOperationInput): Promise<BatchOperationResponse> => {
// ... your custom logic ...
if (someCondition) {
return {
data: result,
shouldHaltBatch: true // This will stop processing remaining URLs
};
}
return { data: result };
};

Controlling Concurrency

Note: We recommend keeping maxWindowsPerSession=1 (the default) for optimal stability and performance. While you can experiment with multiple windows per session, single-window sessions are generally more reliable and easier to manage. Increase maxConcurrentSessions instead if you need more parallelization.

You can control the number of concurrent sessions and windows per session:

const config: BatchOperateConfig = {
maxConcurrentSessions: 30, // Default value is 30
maxWindowsPerSession: 1 // Default value is 1
};

Dynamic URL Addition

Your operation can discover and add new URLs to process during execution. The original batchOperate call will not complete until all operations have completed, including the new additions.

const operation = async (input: BatchOperationInput): Promise<BatchOperationResponse> => {
// ... your custom logic ...
return {
data: result,
additionalUrls: [
{ url: "https://example.com/discovered1" },
{ url: "https://example.com/discovered2" }
]
};
};