Web Crawler Tools — Safe & Structured Web Scraping
What are the Web Crawler Tools?
Langoedge provides powerful web crawling tools designed to retrieve real-time data from external websites. Rather than returning raw, messy HTML, these tools parse pages, strip away advertisements and navigation blocks, and present clean, structured Markdown directly to your AI agents.
The platform exposes one native crawling tool, crawl_url — a deep crawler built on a high-performance engine with built-in sandbox controls and SSRF protection. It needs no account connection and no API key.
1. The Native Web Crawler (crawl_url)
The native web crawler is optimized for speed, deep crawling, and structural output. It supports crawling a single URL or executing a breadth-first search (BFS) across multiple pages.
Security & SSRF Protection
For enterprise security, the native crawler enforces strict Server-Side Request Forgery (SSRF) protections. The gateway verifies every destination IP address before connection:
- Blocked Ranges: All private IP subnets (e.g.,
10.0.0.0/8,192.168.0.0/16,172.16.0.0/12), loopbacks (127.0.0.1,::1), and cloud metadata services (169.254.169.254). - Domain Isolation: Any attempt to query internal workspace microservices or hostnames (like
localhostor cloud dashboard services) will raise anUnsafeURLErrorand halt execution instantly.
Tool Parameters
When configuring a crawl_url action within a graph node, you can define the following settings:
| Parameter | Type | Required | Description |
|---|---|---|---|
url |
string |
Yes | The starting HTTP/HTTPS link to crawl. |
url_patterns |
array[string] |
No | Glob/Regex patterns to match. Only URLs matching these patterns will be crawled. |
allowed_domains |
array[string] |
No | Whitelist of domains. The crawler will not exit these boundaries. |
blocked_domains |
array[string] |
No | Blacklist of domains to skip. |
max_depth |
integer |
No | Levels to crawl beyond the starting page. Defaults to 3. |
max_pages |
integer |
No | Hard cap on pages fetched in one crawl. Defaults to 50. |
include_external |
boolean |
No | Whether to follow links off the starting domain. Defaults to false. |
score_threshold |
float |
No | Minimum relevance score a URL must reach to be crawled. |
scorer_keywords |
array[string] |
No | When set, prioritises pages by keyword relevance — useful for pulling the pricing page out of a large site instead of crawling all of it. |
Technical Output Structure
The tool returns a list of web page content blocks structured as follows:
[
{
"title": "Page Title",
"content": "## Section Heading\n\nThis is structured markdown text extracted from the page.",
"metadata": {
"source": "https://example.com/sub-page"
}
}
]