SkillTotal

Is unclecode/crawl4ai safe?

Some risk - review before installing
Notable — review in context (capabilities are not malware):
  • Python shell/command execution
  • Unsafe deserialization
  • Python dynamic code execution

What to do: Read the findings below before installing: each one opens the exact line of code it was found on.

unclecode/crawl4ai is a PyPI package analyzed by SkillTotal's deterministic static scanner. The scan found no malicious indicators, though 2 risky constructs are reported for review. It can: dynamic code execution, filesystem read, filesystem write, mcp tools detected, network egress and shell execution — capabilities are what the code can do, not a verdict on intent. Risk score 30/100 (medium).

Crawl4AI

python_package · https://github.com/unclecode/crawl4ai
MEDIUM
30
/ 100 risk score
Snapshot · scanned Oct 8, 2026 · Crawl4AI@8afd0a6 · engine 0.56.4 / ruleset 62

Automated static-analysis result. It can contain false positives and false negatives, and is not a claim about the intent of unclecode/crawl4ai's authors. Report a false positive.

Capabilities — what this component can do (not a risk score):
dynamic code executionfilesystem readfilesystem writemcp tools detectednetwork egressshell execution

Behavioral traits

How this component maps to the CSA agentic threat model. Descriptive — it never affects the risk score.

Tool surface
Tool Usage
Execution authority
Tool Access Control / Direct Tool Access
Filesystem reach
Tool Execution Context
Network egress
Interaction & Communication / Direct Communication
Network exposure
Interaction & Communication
Embedded credential access
Tool Execution Context / Agent Service Identity
Unsafe deserialization
General Protections / Input Validation

Findings (10)

HIGHUnsafe deserializationST-DESERIALIZE-PY

It loads data with a format that can rebuild arbitrary objects (e.g. pickle, or unsafe YAML).

Why it matters: Feeding such a loader untrusted data can execute code hidden inside that data.

Fix: Deserialize untrusted data with a safe format/loader: JSON, or yaml.safe_load / Loader=SafeLoader. Reserve pickle/marshal for data you fully control.

HIGHPython dynamic code executionST-DYN-PY

The code turns strings into live code at runtime (eval / new Function / exec).

return Compiler(root).compile(script)

Why it matters: If those strings aren't fixed and trusted, they become a way to run arbitrary code.

Fix: Avoid evaluating dynamically constructed code; if unavoidable, ensure the input is a trusted constant and never derived from external data.

HIGHEnvironment file shipped in the released packagenot scoredST-ENV-SHIPPED

The published artifact contains a .env file. That is where a project keeps its environment secrets; it reaches a release when the packer captures the project root, the same way publisher credentials do.

3 variable(s), values withheld: GROQ_API_KEY=…, OPENAI_API_KEY=…, ANTHROPIC_API_KEY=…

Fix: Treat every value in it as exposed and rotate it, then exclude the file from the release (a `files` allowlist or .npmignore for npm, MANIFEST.in for a Python sdist — .gitignore alone excludes it from neither).

HIGHEmbedded secret / credentialnot scoredST-SECRET-EMBEDDED

A hardcoded credential (API key, token, or private key) is shipped in the code.

- API token of LLM provider <br/> eg: `api_token = "gsk_…[redacted, 54 chars]"`

Why it matters: Anyone who gets the package gets the secret — rotate it and load secrets at runtime instead.

Fix: Remove the secret from the code, rotate it immediately, and load credentials from the environment or a secrets manager at runtime.

HIGHPython shell/command executionST-SHELL-PY

The component can run operating-system commands or spawn processes.

subprocess.check_output(
                            shlex.split(f"lsof -t -i:{self.debugging_port}"),
                            stderr=subprocess.DEVNULL,
                        )
self.browser_process = subprocess.Popen(
                    args, 
                    stdout=subprocess.PIPE, 
                    stderr=subprocess.PIPE,
                    creationflags=subprocess.DETACHED_PROCESS | subprocess.CREATE_N …
self.browser_process = subprocess.Popen(
                    args, 
                    stdout=subprocess.PIPE, 
                    stderr=subprocess.PIPE,
                    preexec_fn=os.setpgrp  # Start in a new process group …
subprocess.run(["taskkill", "/F", "/T", "/PID", str(self.browser_process.pid)])
process = subprocess.run(["tasklist", "/FI", f"PID eq {pid}"], 
                                         capture_output=True, text=True)
subprocess.run(["taskkill", "/F", "/PID", str(pid)], check=True)
subprocess.Popen(browser_cmd + browser_args)
subprocess.check_call(
            [
                sys.executable,
                "-m",
                "playwright",
                "install",
                "--with-deps",
                "--force",
                "chromium", …
subprocess.check_call(
            [
                sys.executable,
                "-m",
                "patchright",
                "install",
                "--with-deps",
                "--force",
                "chromium", …
subprocess.run(
                ["git", "clone", "-b", branch, repo_url, str(repo_folder)],
                stdout=subprocess.DEVNULL,
                stderr=subprocess.DEVNULL,
                check=True,
            )
output = subprocess.check_output(["sysctl", "-n", "hw.memsize"]).decode("utf-8")
xvfb = subprocess.Popen(["Xvfb", ":99", "-screen", "0", "1280x720x24"])
fluxbox = subprocess.Popen(["fluxbox"])
x11vnc = subprocess.Popen(["x11vnc",
                              "-display", ":99",
                              "-nopw", "-forever", "-shared",
                              "-rfbport", "5900", "-quiet"])
novnc = subprocess.Popen(["/opt/novnc/utils/websockify/run",
                              "6080", "localhost:5900",
                              "--web", "/opt/novnc"])
result = subprocess.run(['vm_stat'], capture_output=True, text=True)
result = subprocess.run(
            ["gh", "api", *args],
            capture_output=True, text=True, timeout=30
        )

Why it matters: Powerful and often legitimate — confirm the commands aren't built from untrusted input.

Fix: Confirm the command and its arguments are fully controlled and not derived from untrusted input; avoid shell=True.

MEDIUMServer bound to all network interfacesST-EXPOSE-BIND

A server is bound to all network interfaces (0.0.0.0), not just your own machine.

Why it matters: Without authentication, other hosts on the network can reach it.

Fix: Bind to 127.0.0.1 for local-only use, or require authentication and restrict access if remote exposure is intended.

MEDIUMPython filesystem readST-FS-PY-READ

The component reads files from disk.

with open(path, 'r') as f:
with open(filepath, 'r', encoding='utf-8') as f:
with open(local_file_path, "r", encoding="utf-8") as f:
with open(local_file_path, "r", encoding="utf-8") as f:
with open(cache_path, "r") as f:
with open(cache_path, "r") as f:
return self.index_cache_path.read_text().strip()
with open(self.builtin_config_file, 'r') as f:
with open(config_file) as f:
with open(path) as f:
with open(config_file) as f:
with open(tar_path, "rb") as f:
if marker.exists() and marker.read_text().strip() == value:
self.js_script = (Path(__file__).parent /
                          "script.js").read_text()
with open(f"{home_dir}/schema/organic_schema.json", "r") as f:
with open(f"{home_dir}/schema/top_stories_schema.json", "r") as f:
with open(f"{home_dir}/schema/suggested_query_schema.json", "r") as f:
data = json.loads(path.read_text())
with open(args.filename, "rb") as fp:
with open(script_path, "r") as f:
with open(cache_file_path, "r") as f:
with open(file_path, "r", encoding="utf-8") as f:
with open(cache_file, "r") as f:
with open(self.bm25_index_file, "rb") as f:
with open(file, "r", encoding="utf-8") as f_obj:

Why it matters: Usually legitimate, but worth confirming it can't be steered into reading sensitive files.

Fix: Confirm which files are read and that paths cannot be influenced by untrusted input to reach sensitive locations.

MEDIUMPython filesystem write/deleteST-FS-PY-WRITE

The component writes or deletes files on disk.

with open(path, 'w') as f:
with open(filepath, 'w', encoding='utf-8') as f:
with open(self.log_file, "a", encoding="utf-8") as f:
with open(self.log_file, "a", encoding="utf-8") as f:
with open(cache_path, "w") as f:
self.index_cache_path.write_text(idx)
shutil.rmtree(self.temp_dir)
shutil.rmtree(item) if item.is_dir() else item.unlink()
shutil.rmtree(profile_path)
with open(self.builtin_config_file, 'w') as f:
os.unlink(self.builtin_config_file)
with open(config_file, "w", encoding="utf-8") as f:
with open(output_file, "w", encoding="utf-8") as f:
with open(output_file, "w", encoding="utf-8") as f:
with open(output_file, "w", encoding="utf-8") as f:
with open(output_file, "w", encoding="utf-8") as f:
with open(config_file, "w") as f:
shutil.rmtree(temp_dir, ignore_errors=True)
marker.write_text(value + "\n")
with open(f"{home_dir}/schema/organic_schema.json", "w") as f:
with open(f"{home_dir}/schema/top_stories_schema.json", "w") as f:
with open(f"{home_dir}/schema/suggested_query_schema.json", "w") as f:
path.write_text(json.dumps(data, default=str))
shutil.rmtree(cache_folder)

Why it matters: Usually legitimate, but worth confirming the paths can't be controlled by untrusted input.

Fix: Confirm which files are written/deleted and that paths cannot be influenced by untrusted input.

MEDIUMPython network egressST-NET-PY

The component makes outbound network requests.

response = requests.get(url, params=params, timeout=10)
from aiohttp.client import ClientTimeout
connector = aiohttp.TCPConnector(
                limit=self.max_connections,
                ttl_dns_cache=self.dns_cache_ttl,
                use_dns_cache=True,
                force_close=False
            )
self._session = aiohttp.ClientSession(
                headers=dict(self._BASE_HEADERS),
                connector=connector,
                timeout=ClientTimeout(total=self.DEFAULT_TIMEOUT)
            )
timeout=ClientTimeout(total=self.DEFAULT_TIMEOUT)
timeout = ClientTimeout(
                total=timeout_sec,
                connect=10,
                sock_read=30
            )
self.client = client or httpx.AsyncClient(http2=True, timeout=20, proxy=proxy_url(), headers={
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) +AppleWebKit/537.36 (KHTML, like Gecko) Chrome/123.0.0.0 Safari/537.36" …
async with httpx.AsyncClient(proxy=proxy_url()) as c:
async with aiohttp.ClientSession() as session:
async with session.get(verify_url, timeout=aiohttp.ClientTimeout(total=2)) as response:
async with aiohttp.ClientSession() as session:
async with aiohttp.ClientSession() as session:
self._client = httpx.AsyncClient(
                http2=True,
                timeout=self.timeout,
                follow_redirects=True,
                headers={"User-Agent": self.user_agent}
            )
response = httpx.get(
            f"{test_url}/v1/profiles",
            headers={"X-API-Key": api_key},
            timeout=10.0
        )
response = httpx.post(
                f"{api_url}/v1/profiles",
                headers={"X-API-Key": api_key},
                files={"file": (f"{cloud_name}.tar.gz", f, "application/gzip")},
                data={"name": cloud_name}, …
response = httpx.get(
            f"{api_url}/v1/profiles",
            headers={"X-API-Key": api_key},
            timeout=30.0
        )
response = httpx.get(
            f"{api_url}/v1/profiles",
            headers={"X-API-Key": api_key},
            timeout=30.0
        )

Why it matters: Usually legitimate, but confirm the destinations are expected and no sensitive data leaves.

Fix: Confirm the destination hosts are expected and that no sensitive data is sent off-host.

LOWMCP tool surface detectedST-MCP-DETECTED

An MCP tool surface (manifest or tool definitions) was found.

t.Tool(name=k, description=desc, inputSchema=schema)

Why it matters: Just context — review which tools it offers and their permissions.

Fix: Review the declared MCP tools and their permissions.

How attackers abuse these capabilities

Interactive labs on the attack class behind the rules above. They show the technique, not anything found in unclecode/crawl4ai.

Check your own component

Run the same evidence-backed scan on any MCP server, agent skill, or package.

Scan your own component

How we determine this: deterministic static analysis (regex + AST), evidence-anchored, no code execution. Methodology →