> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.airtop.ai/guides/how-to/batch-operations/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.airtop.ai/_mcp/server. The batch operate SDK helpers allow you to efficiently process multiple URLs in parallel while managing browser sessions and windows automatically. These helpers automatically manage the lifecycle of browser sessions and windows, handling creation and cleanup behind the scenes so you can focus on your own application logic. This guide explains how to use these helpers which are available in both the Node.js and Python Airtop SDKs. ## Python SDK ### Basic Usage ```python from airtop import AsyncAirtop, types, BatchOperationUrl, BatchOperationInput, BatchOperationResponse, BatchOperateConfig from dataclasses import dataclass from typing import Optional, List, Any, Dict async def main(): client = AsyncAirtop(api_key="your-api-key") # Define URLs to process urls = [ BatchOperationUrl(url="https://example.com/page1"), BatchOperationUrl(url="https://example.com/page2"), BatchOperationUrl(url="https://example.com/page3") ] # Define your operation function async def operation(input: BatchOperationInput) -> BatchOperationResponse: # Example: Run a custom query on the page result = await client.windows.page_query( session_id=input.session_id, window_id=input.window_id, prompt="What is the main idea of this page?" ) return BatchOperationResponse( data=result.data.model_response, should_halt_batch=False, additional_urls=[] ) # Execute batch operation results = await client.batch_operate( urls=urls, operation=operation, ) # Results will be an array containing the responses from each operation # Example: # [ # "The main idea of this page is about cloud computing services and their benefits for enterprise...", # "This page discusses machine learning algorithms and their applications in data science...", # "The page focuses on cybersecurity best practices for organizations..." # ] ``` ### Advanced Features #### Error Handling By default, the Batch Operate helpers will *not* automatically retry failed operations. However, you can implement this or other custom error handling by providing an `on_error` callback. ```python async def handle_error(error: BatchOperationError): print(f"Error occurred: {error.error}") print(f"Session ID: {error.session_id}") print(f"Window ID: {error.window_id}") print(f"URLs affected: {error.operation_urls}") config = BatchOperateConfig( on_error=handle_error ) ``` #### Session Configuration You can customize the session configuration for all sessions created during the batch operation. A very helpful option is the `profile_name` parameter. This allows you to re-use a browser [Profile](/guides/how-to/saving-a-profile) that has sign-in credentials and cookies for websites that require authentication. ```python session = await client.sessions.create() # ... Your custom logic, i.e. authentication ... # Terminate the session to persist the profile await client.sessions.terminate(session.data.id) config = BatchOperateConfig( session_config=types.SessionConfigV1( timeout_minutes=10, # Terminate the session after 10 mins of inactivity profile_name="my-linkedin-profile" # Use a profile that has been signed in to LinkedIn ) ) results = await client.batch_operate( urls=urls, operation=operation, config=config ) ``` #### Halting Operations You can stop the batch processing at any point by returning `should_halt_batch=True`: ```python async def operation(input: BatchOperationInput) -> BatchOperationResponse: # ... your custom logic ... if some_condition: return BatchOperationResponse( data=result, should_halt_batch=True # This will stop processing remaining URLs ) return BatchOperationResponse(data=result) ``` #### Controlling Concurrency > **Note**: We recommend keeping `max_windows_per_session=1` (the default) for optimal stability and performance. While you can experiment with multiple windows per session, single-window sessions are generally more reliable and easier to manage. Increase `max_concurrent_sessions` instead if you need more parallelization. You can control the number of concurrent sessions and windows per session: ```python config = BatchOperateConfig( max_concurrent_sessions=30, # Default value is 30 max_windows_per_session=1 # Default value is 1 ) ``` #### Dynamic URL Addition Your operation can discover and add new URLs to process during execution. The original batch\_operate call will not complete until all operations have completed, including the new additions. ```python async def operation(input: BatchOperationInput) -> BatchOperationResponse: # ... your scraping logic ... new_urls = [ BatchOperationUrl(url="https://example.com/discovered1"), BatchOperationUrl(url="https://example.com/discovered2") ] return BatchOperationResponse( data=result, additional_urls=new_urls ) ``` ## Node.js SDK ### Basic Usage ```typescript import { Airtop, BatchOperationUrl, BatchOperationInput, BatchOperationResponse, BatchOperateConfig } from '@airtop/sdk'; async function main() { const client = new Airtop({ apiKey: "your-api-key" }); // Define URLs to process const urls: BatchOperationUrl[] = [ { url: "https://example.com/page1" }, { url: "https://example.com/page2" }, { url: "https://example.com/page3" } ]; // Define your operation function const operation = async (input: BatchOperationInput): Promise => { const { windowId, sessionId } = input; // Example: Run a custom query on the page const result = await client.windows.pageQuery({ sessionId, windowId, prompt: "What is the main idea of this page?" }); return { data: result.data.modelResponse, shouldHaltBatch: false, additionalUrls: [] }; }; // Execute batch operation const results = await client.batchOperate({ urls, operation }); // Results will be an array containing the responses from each operation // Example results: // [ // "This page is about cloud computing services and infrastructure", // "The main topic is machine learning and AI applications", // "This page discusses data analytics and visualization tools" // ] } ``` ### Advanced Features #### Error Handling By default, the Batch Operate helpers will *not* automatically retry failed operations. However, you can implement this or other custom error handling by providing an `onError` callback. ```typescript const config: BatchOperateConfig = { onError: async (error: BatchOperationError) => { console.error(`Error occurred: ${error.error}`); console.error(`Session ID: ${error.sessionId}`); console.error(`Window ID: ${error.windowId}`); console.error(`URLs affected:`, error.operationUrls); } }; ``` #### Session Configuration You can customize the session configuration for all sessions created during the batch operation. A very helpful option is the `profileName` parameter. This allows you to re-use a browser [Profile](/guides/how-to/saving-a-profile) that has sign-in credentials and cookies for websites that require authentication. ```typescript // Create and setup initial session const session = await client.sessions.create(); // ... Your custom logic, i.e. authentication ... // Terminate the session to persist the profile await client.sessions.terminate({ sessionId: session.data.id }); const config: BatchOperateConfig = { sessionConfig: { timeoutMinutes: 10, // Terminate the session after 10 mins of inactivity profileName: "my-linkedin-profile", // Use a profile that has been signed in to LinkedIn } }; const results = await client.batchOperate({ urls, operation, config }); ``` #### Halting Operations You can stop the batch processing at any point by returning `shouldHaltBatch: true`: ```typescript const operation = async (input: BatchOperationInput): Promise => { // ... your custom logic ... if (someCondition) { return { data: result, shouldHaltBatch: true // This will stop processing remaining URLs }; } return { data: result }; }; ``` #### Controlling Concurrency > **Note**: We recommend keeping `maxWindowsPerSession=1` (the default) for optimal stability and performance. While you can experiment with multiple windows per session, single-window sessions are generally more reliable and easier to manage. Increase `maxConcurrentSessions` instead if you need more parallelization. You can control the number of concurrent sessions and windows per session: ```typescript const config: BatchOperateConfig = { maxConcurrentSessions: 30, // Default value is 30 maxWindowsPerSession: 1 // Default value is 1 }; ``` #### Dynamic URL Addition Your operation can discover and add new URLs to process during execution. The original batchOperate call will not complete until all operations have completed, including the new additions. ```typescript const operation = async (input: BatchOperationInput): Promise => { // ... your custom logic ... return { data: result, additionalUrls: [ { url: "https://example.com/discovered1" }, { url: "https://example.com/discovered2" } ] }; }; ```