About this workflow
Search & Summarize Web Data with Perplexity, Gemini AI & Bright Data to Webhooks – Workflow Analysis
This N8n workflow template automates the process of searching for information on the web using Perplexity via Bright Data’s Web Scraper API, extracting readable content from HTML responses using Google Gemini AI, and summarizing the extracted data through a LangChain-powered summarization chain. The final summarized output is then sent to an external service via a webhook notification. Designed for AI-enhanced data extraction and processing, this workflow demonstrates advanced integration between web scraping, language models, and conditional logic within the n8n automation platform.
What the Workflow Does
The workflow initiates a search query on Perplexity.ai (via Bright Data's dataset triggering system), waits until the snapshot of results is ready, downloads it, extracts clean, human-readable text from the HTML response using Google Gemini, summarizes the content using another Gemini model, and finally sends the summary to a configurable webhook endpoint. It includes error handling, polling mechanisms, and structured AI processing to ensure robustness and accuracy in automated research tasks.
It effectively turns unstructured web content into concise, actionable summaries—ideal for competitive intelligence, market research, or knowledge aggregation systems.
Key Features and Capabilities
- Automated Web Research: Triggers real-time searches on Perplexity using Bright Data’s API.
- AI-Powered Content Extraction: Uses Google Gemini to extract only the relevant "readable content" from complex HTML.
- Document Summarization: Leverages LangChain’s summarization chain with Gemini Flash models to generate condensed insights.
- Polling Mechanism: Implements a loop that checks job status until the scraper returns a completed snapshot.
- Error Handling: Evaluates whether errors occurred during scraping before proceeding.
- Webhook Integration: Outputs final results to any external system via HTTP POST.
- Modular AI Node Usage: Demonstrates use of
@n8n/nodes-langchainnodes for LLM chaining, document loading, and text splitting. - Credential Management: Securely uses stored credentials for both Bright Data and Google Gemini APIs.
Main Nodes and Their Purposes
| Node Name | Type | Purpose |
|---|---|---|
| When clicking ‘Test workflow’ | Manual Trigger | Starts the workflow manually for testing purposes. |
| Perplexity Search Request | HTTP Request | Sends a POST request to Bright Data’s API to trigger a search on Perplexity.ai with a predefined prompt (tell me about BrightData). |
| Set Snapshot Id | Set | Stores the returned snapshot_id from the initial trigger for later use in polling and downloading. |
| Check Snapshot Status | HTTP Request | Polls the Bright Data API to check if the snapshot generation is complete by querying progress. |
| If (status check) | If Condition | Checks whether the status field in the response equals "ready"; routes execution accordingly. |
| Wait | Wait | Pauses execution for 30 seconds before retrying the status check if not ready yet. |
| Check on the errors | If Condition | Validates whether the errors count is zero before proceeding to download. |
| Download Snapshot | HTTP Request | Retrieves the full JSON result of the completed snapshot using the snapshot_id. |
| Readable Data Extractor | Information Extractor (LangChain) | Uses Google Gemini to parse answer_html and extract only the "readable content" as plain text. |
| Google Gemini Chat Model1 | LLM Node | Provides the language model backend for the data extractor (uses gemini-2.0-flash-exp). |
| Summarization of search result | Summarization Chain (LangChain) | Takes the cleaned text and generates a concise summary using a document loader pattern. |
| Default Data Loader | Document Loader (LangChain) | Feeds processed documents into the summarization chain. |
| Recursive Character Text Splitter | Text Splitter (LangChain) | Prepares long texts by splitting them into manageable chunks (with 100-character overlap). |
| Google Gemini Chat Model | LLM Node | Supplies the LLM engine for the summarization chain (uses experimental gemini-2.0-flash-thinking-exp-01-21 model). |
| Webhook Notifier | HTTP Request | Sends the final output (summary) to an external URL (webhook.site placeholder). |
| Sticky Notes | Annotation | Provides documentation and guidance within the canvas for users setting up credentials or understanding flow logic. |
Use Cases and Benefits
Use Cases:
- Market Intelligence Gathering: Automatically summarize competitor analysis, product reviews, or industry trends.
- Content Aggregation Systems: Build internal knowledge bases by pulling and condensing online articles or Q&A platforms.
- Research Automation: Replace manual browsing and note-taking with AI-driven search and summarization pipelines.
- Lead Generation & Outreach Prep: Quickly gather background info on companies or individuals before outreach.
- News Monitoring: Track mentions of brands, technologies, or events across sources like Perplexity-curated answers.
Benefits:
- Time Efficiency: Eliminates hours of manual reading and summarizing.
- Accuracy: AI ensures only meaningful content is retained and summarized.
- Scalability: Can be scheduled or triggered in bulk for multiple queries.
- Integration-Friendly: Output delivered via webhook enables connection to Slack, CRM, databases, or email engines.
- Error Resilience: Built-in loops and conditionals prevent failures due to timing issues or transient errors.
Step-by-Step Workflow Logic
-
Trigger Execution
- User clicks “Test workflow” → activates the manual trigger node.
-
Initiate Search Query
- A POST request is sent to
https://api.brightdata.com/datasets/v3/triggerwith:- Target URL:
https://www.perplexity.ai - Prompt:
"tell me about BrightData" - Country:
"US" - Dataset ID and auth headers included.
- Target URL:
- Response contains a
snapshot_id.
- A POST request is sent to
-
Store Snapshot ID
- The
Set Snapshot Idnode saves thesnapshot_idfrom the previous response for reuse.
- The
-
Poll for Completion
Check Snapshot Statusrequests progress at.../progress/{snapshot_id}.- An If node checks if
status === "ready":- If true → proceeds.
- If false → triggers Wait (30s) → loops back to re-check status.
-
Error Validation
- Once status is ready, another If node checks if
errors.toString() === "0". - Ensures no errors occurred during scraping before continuing.
- Once status is ready, another If node checks if
-
Download Final Results
- On successful validation, the snapshot is downloaded in JSON format from
.../snapshot/{snapshot_id}.
- On successful validation, the snapshot is downloaded in JSON format from
-
Extract Readable Content
- The
answer_htmlfield is passed to the Readable Data Extractor. - Powered by Google Gemini Flash Exp, it isolates clean, readable text from noisy HTML.
- The
-
Prepare for Summarization
- The extracted text is fed into the Summarization Chain.
- Before summarizing, it passes through:
- Recursive Character Text Splitter → breaks large text into chunks.
- Default Data Loader → loads split text as AI-ready documents.
-
Generate Summary
- The Summarization of search result node uses the second Gemini model (
thinking-exp) to create a coherent, high-level summary.
- The Summarization of search result node uses the second Gemini model (
-
Send Output via Webhook
- Final summary (
$json.output) is sent via POST tohttps://webhook.site/.... - Can be replaced with any desired endpoint (e.g., Slack, Airtable, custom API).
- Final summary (
This workflow exemplifies how n8n combines low-code automation with powerful AI capabilities to transform raw web data into intelligent outputs—making it ideal for developers, researchers, and business analysts looking to scale their information-processing workflows.
How to use: download the JSON, then in n8n choose “Import from File”.