Is D4Vinci/Scrapling safe?
- Python shell/command execution
- Unsafe deserialization
- Python network egress
What to do: Read the findings below before installing: each one opens the exact line of code it was found on.
D4Vinci/Scrapling is a PyPI package analyzed by SkillTotal's deterministic static scanner. The scan found no malicious indicators, though 2 risky constructs are reported for review. It can: filesystem read, filesystem write, mcp tools detected, network egress and shell execution — capabilities are what the code can do, not a verdict on intent. Risk score 30/100 (medium).
scrapling 0.4.15
Automated static-analysis result. It can contain false positives and false negatives, and is not a claim about the intent of D4Vinci/Scrapling's authors. Report a false positive.
Behavioral traits
How this component maps to the CSA agentic threat model. Descriptive — it never affects the risk score.
Findings (7)
It loads data with a format that can rebuild arbitrary objects (e.g. pickle, or unsafe YAML).
data: CheckpointData = pickle.loads(content)
Why it matters: Feeding such a loader untrusted data can execute code hidden inside that data.
Fix: Deserialize untrusted data with a safe format/loader: JSON, or yaml.safe_load / Loader=SafeLoader. Reserve pickle/marshal for data you fully control.
The component can run operating-system commands or spawn processes.
_ = check_output(cmd, shell=False) # nosec B603
Why it matters: Powerful and often legitimate — confirm the commands aren't built from untrusted input.
Fix: Confirm the command and its arguments are fully controlled and not derived from untrusted input; avoid shell=True.
A server is bound to all network interfaces (0.0.0.0), not just your own machine.
help="The host to use if streamable-http transport is enabled. Pass '0.0.0.0' to accept connections from "
Why it matters: Without authentication, other hosts on the network can reach it.
Fix: Bind to 127.0.0.1 for local-only use, or require authentication and restrict access if remote exposure is intended.
The component reads files from disk.
browsers = loads(path.read_bytes()).get("browsers", [])Why it matters: Usually legitimate, but worth confirming it can't be steered into reading sensitive files.
Fix: Confirm which files are read and that paths cannot be influenced by untrusted input to reach sensitive locations.
The component writes or deletes files on disk.
shutil.rmtree(path)
with open(fd, "w", encoding=page.encoding) as f:
with open(filename, "w", encoding=page.encoding) as f:
file.write_bytes(orjson.dumps(list(self), option=options))
with open(path, "wb") as f:
with open(file, "w", newline="", encoding="utf-8") as f:
await file.write_text(item["markdown"], encoding="utf-8")
Why it matters: Usually legitimate, but worth confirming the paths can't be controlled by untrusted input.
Fix: Confirm which files are written/deleted and that paths cannot be influenced by untrusted input.
The component makes outbound network requests.
import requests
req = requests.get("https://books.toscrape.com/index.html")Why it matters: Usually legitimate, but confirm the destinations are expected and no sensitive data leaves.
Fix: Confirm the destination hosts are expected and that no sensitive data is sent off-host.
An MCP tool surface (manifest or tool definitions) was found.
server.add_tool(
Why it matters: Just context — review which tools it offers and their permissions.
Fix: Review the declared MCP tools and their permissions.
How attackers abuse these capabilities
Interactive labs on the attack class behind the rules above. They show the technique, not anything found in D4Vinci/Scrapling.
Check your own component
Run the same evidence-backed scan on any MCP server, agent skill, or package.
Scan your own componentHow we determine this: deterministic static analysis (regex + AST), evidence-anchored, no code execution. Methodology →