Is infiniflow/ragflow safe?
- Python shell/command execution
- Environment file shipped in the released package
- Embedded secret / credential
ragflow is an AI python_package analyzed by SkillTotal's deterministic static scanner. The scan found no malicious indicators, though 9 risky constructs are reported for review. It can: filesystem read, filesystem write, mcp tools detected, network egress, scoped identity and shell execution — capabilities are what the code can do, not a verdict on intent. Risk score 0/100 (low).
ragflow 1.0.0-rc1
Automated static-analysis result. It can contain false positives and false negatives, and is not a claim about the intent of infiniflow/ragflow's authors. Report a false positive.
Behavioral traits
How this component maps to the CSA agentic threat model. Descriptive — it never affects the risk score.
Findings (9)
The published artifact contains a .env file. That is where a project keeps its environment secrets; it reaches a release when the packer captures the project root, the same way publisher credentials do.
119 variable(s), values withheld: DOC_ENGINE=…, DB_TYPE=…, DEVICE=…, METADATA_DB_PROFILE=…, COMPOSE_PROFILES=…, STACK_VERSION=…, ES_HOST=…, ES_PORT=…, and 111 more
6 variable(s), values withheld: MINIO_HOST=…, MINIO_USER=…, MINIO_PASSWORD=…, MINIO_BUCKET=…, MINIO_PREFIX_PATH=…, STORAGE_IMPL=…
3 variable(s), values withheld: PORT=…, DID_YOU_KNOW=…, VITE_DEFAULT_LANGUAGE_CODE=…
1 variable(s), values withheld: VITE_BASE_URL=…
2 variable(s), values withheld: VITE_BASE_URL=…, VITE_BUILD_SOURCEMAP=…
Fix: Treat every value in it as exposed and rotate it, then exclude the file from the release (a `files` allowlist or .npmignore for npm, MANIFEST.in for a Python sdist — .gitignore alone excludes it from neither).
A hardcoded credential (API key, token, or private key) is shipped in the code.
password: 'infi…[redacted, 21 chars]'
Why it matters: Anyone who gets the package gets the secret — rotate it and load secrets at runtime instead.
Fix: Remove the secret from the code, rotate it immediately, and load credentials from the environment or a secrets manager at runtime.
The component can run operating-system commands or spawn processes.
proc = await asyncio.create_subprocess_exec(
npm,
"install",
"--no-fund",
"--no-audit",
cwd=str(gateway_dir),
)self._process = await asyncio.create_subprocess_exec(
*cfg.command,
cwd=cfg.cwd,
env=env,
)commit = subprocess.run(
["git", "rev-parse", "HEAD"],
cwd=REPO_ROOT,
capture_output=True,
text=True,
check=True,
).stdout.strip()result = subprocess.check_output(["tput", "colors"], stderr=subprocess.DEVNULL, text=True)
result = subprocess.run(cmd, check=False)
proc = subprocess.run(
["git", *args, "-z"],
check=True,
capture_output=True,
)Why it matters: Powerful and often legitimate — confirm the commands aren't built from untrusted input.
Fix: Confirm the command and its arguments are fully controlled and not derived from untrusted input; avoid shell=True.
The component reads files from disk.
case = json.loads(path.read_text(encoding="utf-8"))
existing = target.read_text(encoding="utf-8") if target.exists() else None
sections = parser()(str(sample_path), sample_path.read_bytes(), TOKEN_SIZE, DELIMITER, True)
case = json.loads(path.read_text(encoding="utf-8"))
with open(path, "r", encoding="utf-8") as f:
with open(path, "rb") as f:
with open(path, "rb") as f:
with open(corpus_path, encoding="utf-8") as fh:
with open(args.catalog, encoding="utf-8") as handle:
with open(progress_file, "r") as f:
return path.read_bytes()
json.loads(path.read_text(encoding="utf-8"))
yaml.safe_load(path.read_text(encoding="utf-8"))
data = _read_bytes(path)
data = _read_bytes(path)
data = _read_bytes(path)
data = _read_bytes(path)
source = path.read_text()
with open(os.path.join(args.out_dir, "quads.json")) as f:
with open(os.path.join(args.out_dir, "quads_meta.json")) as f:
with open(meta_path) as f:
Why it matters: Usually legitimate, but worth confirming it can't be steered into reading sensitive files.
Fix: Confirm which files are read and that paths cannot be influenced by untrusted input to reach sensitive locations.
The component writes or deletes files on disk.
with open(src_b64, "w") as f:
with open(exp_b64, "w") as f:
with open(meta_path, "w") as f:
MANIFEST.write_text(
json.dumps(
{
"captured_at": datetime.now(timezone.utc).isoformat(timespec="seconds"),
"commit": commit,
"python": sys.version.split()[0], …target.write_text(first, encoding="utf-8")
path.write_text(json.dumps(result, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
shutil.rmtree(version_dir)
shutil.rmtree(stale)
shutil.rmtree(normalized)
shutil.rmtree(stale)
shutil.rmtree(normalized)
with open(filename, "wb") as file:
shutil.rmtree(version_dir)
with open(args.out, "w", encoding="utf-8") as fh:
with open(output, "w") as f:
with open(progress_file, "w") as f:
path.write_text(new_text, encoding="utf-8")
path.write_bytes(data.replace(b"\r\n", b"\n"))
with open(os.path.join(args.out_dir, "quads.json"), "w") as f:
with open(os.path.join(args.out_dir, "quads_meta.json"), "w") as f:
Why it matters: Usually legitimate, but worth confirming the paths can't be controlled by untrusted input.
Fix: Confirm which files are written/deleted and that paths cannot be influenced by untrusted input.
The component makes outbound network requests.
import http from 'node:http';
import axios from 'axios';
const ret = await axios.get(api, {const { data } = await axios.get(url, { headers: httpHeaders });item.promise = fetch(url, { headers: { [Authorization]: authorization } })const response = await fetch(api.chatsTranscriptions, {[ProgrammingLanguage.Javascript]: `const axios = require('axios');const response = await axios.get('https://github.com/infiniflow/ragflow');import axios from 'axios';
const ret = await axios.get('/conf.json');const response = await fetch(url, {const response = await fetch(url, {const response = await fetch(url, {import { type AxiosResponseHeaders } from 'axios';const response = await fetch(url, { headers });const response = await fetch(downloadUrl);
const response = await fetch('/api/v1/skills/search', {const indexResponse = await fetch('/api/v1/skills/index', {const response = await fetch('/api/v1/skills/status', {* fetch (mount auto-fetch and manual "List models" clicks). Used by
// Every catalog fetch (mount auto-fetch and "List models" clicks)
* - the API fetch (with edit-mode pre-check seeding) and modal-reset effect
import axios from 'axios';
const request = axios.create({* deliberately uses the native `fetch` API instead of the shared axios instance
Why it matters: Usually legitimate, but confirm the destinations are expected and no sensitive data leaves.
Fix: Confirm the destination hosts are expected and that no sensitive data is sent off-host.
The component makes outbound network requests.
import aiohttp
timeout = aiohttp.ClientTimeout(total=max(int(self.account.timeout_secs), 1))
self._http_session = aiohttp.ClientSession(timeout=timeout)
import urllib.request
opener = urllib.request.build_opener()
urllib.request.install_opener(opener)
with urllib.request.urlopen(sidecar_url) as resp:
urllib.request.urlretrieve(download_url, filename)
import requests
response = requests.get(url, stream=True)
resp = requests.get(sidecar_url, timeout=30)
import requests
res = requests.post(url=self.api_url + path, json=json, headers=self.authorization_header, stream=stream, files=files)
res = requests.get(url=self.api_url + path, params=params, headers=self.authorization_header, json=json)
res = requests.delete(url=self.api_url + path, json=json, headers=self.authorization_header)
res = requests.put(url=self.api_url + path, json=json, headers=self.authorization_header)
res = requests.patch(url=self.api_url + path, json=json, headers=self.authorization_header)
import requests
response = requests.get(url_new_conversation, headers=headers, params=params_new_conversation)
response = requests.post(url_completion, headers=headers, json=payload_completion)
import aiohttp
timeout = aiohttp.ClientTimeout(total=self.config.timeout)
self.session = aiohttp.ClientSession(headers=headers, timeout=timeout)
Why it matters: Usually legitimate, but confirm the destinations are expected and no sensitive data leaves.
Fix: Confirm the destination hosts are expected and that no sensitive data is sent off-host.
A short-lived, scoped, assumed identity was detected — an STS AssumeRole / session token, a cloud managed or workload identity, an impersonated service account, a projected Kubernetes service-account token, or a dynamic-secret broker. Tools authenticate with a narrowly-scoped credential that expires, rather than a long-lived embedded service credential. (7 occurrence(s) shown as evidence).
{ label: t('setting.dataSourceOptionAssumeRole'), value: 'assume_role' },return authMode === 'assume_role' && bucketType === 's3';
| 'assume_role'
'assume_role',
assume_role: [],
value: 'assume_role',
{authMode === 'assume_role' && (Fix: A scoped, short-lived identity is the smallest-blast-radius execution context. Confirm the assumed role / requested scope grants only the permissions the tool needs, and that the token lifetime is minimal.
An MCP tool surface (manifest or tool definitions) was found.
// unreachable too, and structured_retrieve has neither a @tool decorator nor a caller.
Why it matters: Just context — review which tools it offers and their permissions.
Fix: Review the declared MCP tools and their permissions.
Check your own component
Run the same evidence-backed scan on any MCP server, agent skill, or package.
Scan your own componentHow we determine this: deterministic static analysis (regex + AST), evidence-anchored, no code execution. Methodology →