← Back to all guides

Web Crawler Tools — Safe & Structured Web Scraping

Langoedge Team3 min read

What are the Web Crawler Tools?

Langoedge provides powerful web crawling tools designed to retrieve real-time data from external websites. Rather than returning raw, messy HTML, these tools parse pages, strip away advertisements and navigation blocks, and present clean, structured Markdown directly to your AI agents.

The platform exposes one native crawling tool, crawl_url — a deep crawler built on a high-performance engine with built-in sandbox controls and SSRF protection. It needs no account connection and no API key.


1. The Native Web Crawler (crawl_url)

The native web crawler is optimized for speed, deep crawling, and structural output. It supports crawling a single URL or executing a breadth-first search (BFS) across multiple pages.

Security & SSRF Protection

For enterprise security, the native crawler enforces strict Server-Side Request Forgery (SSRF) protections. The gateway verifies every destination IP address before connection:

  • Blocked Ranges: All private IP subnets (e.g., 10.0.0.0/8, 192.168.0.0/16, 172.16.0.0/12), loopbacks (127.0.0.1, ::1), and cloud metadata services (169.254.169.254).
  • Domain Isolation: Any attempt to query internal workspace microservices or hostnames (like localhost or cloud dashboard services) will raise an UnsafeURLError and halt execution instantly.

Tool Parameters

When configuring a crawl_url action within a graph node, you can define the following settings:

Parameter Type Required Description
url string Yes The starting HTTP/HTTPS link to crawl.
url_patterns array[string] No Glob/Regex patterns to match. Only URLs matching these patterns will be crawled.
allowed_domains array[string] No Whitelist of domains. The crawler will not exit these boundaries.
blocked_domains array[string] No Blacklist of domains to skip.
max_depth integer No Levels to crawl beyond the starting page. Defaults to 3.
max_pages integer No Hard cap on pages fetched in one crawl. Defaults to 50.
include_external boolean No Whether to follow links off the starting domain. Defaults to false.
score_threshold float No Minimum relevance score a URL must reach to be crawled.
scorer_keywords array[string] No When set, prioritises pages by keyword relevance — useful for pulling the pricing page out of a large site instead of crawling all of it.

Technical Output Structure

The tool returns a list of web page content blocks structured as follows:

[
  {
    "title": "Page Title",
    "content": "## Section Heading\n\nThis is structured markdown text extracted from the page.",
    "metadata": {
      "source": "https://example.com/sub-page"
    }
  }
]

Frequently Asked Questions

Does crawl_url support JavaScript-rendered pages?
Yes. The crawler uses a headless browser configuration to render pages before extracting text. There is no separate Firecrawl tool — earlier documentation described a `fire_crawl_url` fallback that was never part of the shipped platform. If a site blocks the crawler outright, fetch the content another way, such as an API tool against the site's own endpoint.
How deep does the native crawler go?
By default it follows links **3 levels** deep and stops at **50 pages**, staying on the starting domain unless `include_external` is set. Both limits are adjustable per call via `max_depth` and `max_pages`.
Can I crawl intranet pages?
No. Due to the SSRF firewalls protecting the Langoedge deployment cluster, the crawler cannot connect to local network addresses or non-public domains.
LT

Langoedge Team

The Langoedge engineering team builds AI agent infrastructure that empowers businesses to deploy reliable, observable AI staff. Follow Langoedge Team on LinkedIn for product updates and architectural deep dives.