Is tokenizers safe?
- Python filesystem read
- Python network egress
- Python filesystem write/delete
tokenizers is an AI python_package analyzed by SkillTotal's deterministic static scanner. The scan found no malicious indicators, though 3 risky constructs are reported for review. It can: filesystem read, filesystem write and network egress — capabilities are what the code can do, not a verdict on intent. Risk score 0/100 (low).
tokenizers 0.23.1
Automated static-analysis result. It can contain false positives and false negatives, and is not a claim about the intent of tokenizers's authors. Report a false positive.
Behavioral traits
How this component maps to the CSA agentic threat model. Descriptive — it never affects the risk score.
Findings (3)
The component reads files from disk.
m.ParseFromString(open(filename, "rb").read())
with open(filename, "r") as f:
with open(self._model, "r") as model_f:
with open(args.input_file, "r") as f:
with open(args.input_file, "r", encoding="utf-8-sig") as f:
m.ParseFromString(open(filename, "rb").read())
with open(css_filename) as f:
Why it matters: Usually legitimate, but worth confirming it can't be steered into reading sensitive files.
Fix: Confirm which files are read and that paths cannot be influenced by untrusted input to reach sensitive locations.
The component writes or deletes files on disk.
with open(args.vocab_output_path, "w") as vocab_f:
with open(args.merges_output_path, "w") as merges_f:
remove(args.model)
os.remove(f"{args.model_prefix}.model")os.remove(f"{args.model_prefix}.vocab")Why it matters: Usually legitimate, but worth confirming the paths can't be controlled by untrusted input.
Fix: Confirm which files are written/deleted and that paths cannot be influenced by untrusted input.
The component makes outbound network requests.
from requests import get
response = get(args.model, allow_redirects=True)
Why it matters: Usually legitimate, but confirm the destinations are expected and no sensitive data leaves.
Fix: Confirm the destination hosts are expected and that no sensitive data is sent off-host.
Check your own component
Run the same evidence-backed scan on any MCP server, agent skill, or package.
Scan your own componentHow we determine this: deterministic static analysis (regex + AST), evidence-anchored, no code execution. Methodology →