llm-abuse · agent-harness · deepseek-harness · api-key-abuse · mcp · fofa · bug-bounty-automation · model-endpoint
DeepSeek Harness Abuse: An Autonomous Agent Exposes Its Operator's Host
This is an operator stealing inference from an internet-exposed OpenAI-compatible model endpoint by pointing the open-source DeepSeek Harness coding agent at it, whose auto-approving tool loop then ran the endpoint's tool calls on the operator's own Windows workstation and posted the output back.
By Davis Zheng·
TLP:CLEAR. Cleared for public release. Captured by the Kinryū Labs honeypot sensor network. Indicators below are defanged.
Follow-up to our earlier report DeepSeek Harness Against a Stolen LLM Gateway: Tool Calls Run on the Caller.
Executive summary
- 492,150characters in one posted prompt
- 86distinct command outputs from the caller's host
- 11agent turns in five minutes
- 3hosts running the same agent harness
Three addresses in China and Hong Kong drove an internet-exposed OpenAI-compatible model endpoint between 20 and 23 September 2026, using the open-source DeepSeek Harness coding agent under the user agent deepseek-harness/0.1.5-rc.2. The two Hong Kong addresses share an API key; the third shares only the public user agent. On 23 September an eleven-turn session from 182.91.103[.]32 ran the endpoint’s tool calls on its own Windows workstation and posted the output back.
The harness approved tool calls automatically, so the endpoint read files and directory listings from the operator’s workstation. The returned config.toml sets permission_mode = "always-approve" with yolo = false, so every tool call the remote endpoint emitted ran without a human confirming it. When a harness in that mode points at an endpoint the operator does not control, the endpoint issues commands that run on the client host.
If you run an OpenAI-compatible model endpoint, verify bearer tokens on /v1/chat/completions and alert on any key that is absent from your issued-key register. If you run an agent harness, keep auto-approval away from any model endpoint you do not control.
Key judgments
- An operator pointed the open-source DeepSeek Harness agent at an exposed OpenAI-compatible endpoint and its client executed the endpoint's tool calls on its own machine, returning the output. Across three sessions from 182.91.103[.]32 there are 279 tool-result postings, 86 of them distinct across 92 distinct tool-call identifiers, whose bodies are PowerShell object listings and file contents rooted at C:\Users\cheng. The same behaviour appears in smaller volume from 203.175.15[.]28 and 103.85.74[.]25, with seven and five result postings respectively. High confidence.
- The returned config.toml sets permission_mode = "always-approve" with yolo = false, so the harness ran every tool call without a human confirming it; that auto-approval is what turned the operator's workstation into the thing being collected. One of the returned files is the operator's own config.toml, which sets permission_mode = "always-approve" with yolo = false, so every tool call the remote endpoint emitted was executed without a human confirming it. High confidence.
- The operator is most likely a Chinese-speaking vulnerability hunter automating bug-bounty and source-audit work. The returned AGENTS.md declares the agent's identity as an authorised researcher doing black-box SRC hunting and white-box 0day auditing, forbids it from asking whether to continue, and its config registers a FOFA asset-search MCP server and a Playwright browser MCP. Moderate confidence.
- The two Hong Kong addresses share credential material and are linked on that basis; the China Unicom address shares only the public harness user agent, so whether it is the same operator is not established. The harness user agent appears from exactly three sources in three /24s over activity from 20 to 23 September 2026. Low confidence.
- We read those throwaway strings as the operator testing whether the endpoint enforced authentication at all before committing the full harness to it; the observed sequence is the three strings, then the long session. Before the long session the same address submitted the throwaway keys 141414, 41414141 and sk-14141242…, and in one earlier session put the string sk-1234 in the model field rather than the key field. It then proceeded to a five-minute eleven-turn agent session. Moderate confidence.
Timeline
- 2026-09-20 12:22 UTC203.175.15[.]28 enumerates models with python-requests/2.28.1, already authenticated with the shared API key; the key was issued before this request
- 2026-09-20 12:34 UTC103.85.74[.]25, in the same Hong Kong hosting organisation, presents the identical key
- 2026-09-21 11:26 UTCFirst appearance of the deepseek-harness/0.1.5-rc.2 user agent across the address set
- 2026-09-23 17:40 UTC182.91.103[.]32 sends two short Chinese liveness prompts under a different client with a throwaway key
- 2026-09-23 18:05 UTCFirst DeepSeek Harness session from 182.91.103[.]32; the client begins returning locally executed tool results
- 2026-09-23 18:23 UTCThe eleven-turn session opens, its first prompt 44,402 characters long
- 2026-09-23 18:23 UTCThe client posts back a PowerShell listing of its own workspace
- 2026-09-23 18:24 UTCThe client posts back the full text of its operating rule file and, shortly after, its harness configuration
- 2026-09-23 18:27 UTCTurn 10 switches the model field from glm-5.2 to deepseek-chat, then back on turn 11
- 2026-09-23 18:28 UTCEleventh and final turn of the session; no further contact from the address in the logged activity
The session, in order
12:22:32 GET /v1/models User-Agent: python-requests/2.28.1 Authorization: Bearer sk-0af475b8…
17:40:10 [user] 图片背景是什么颜色
18:05:08 [system] You are an AI agent powered by DeepSeek Harness. You are a coding agent powered by the glm-5.2 model.
18:23:39 FullName Length Mode C:\Users\cheng\Desktop\clown-src-6k-skill\bin d----- C:\Users\cheng\Desktop\clown-src-6k-skill\mcp-servers d----- C:\Users\cheng\Desktop\clown-src-6k-skill\AGENTS.md 5517 ------ C:\Users\cheng\Desktop\clown-src-6k-skill\config.toml 1044 ------ C:\Users\cheng\Desktop\clown-src-6k-skill\给朋友的提示词.txt 4370 ------
18:24:21 C:\Users\cheng\.dsh --- Name ---- dsh-purge node_modules plugins profiles sessions staging storages .anonymous-user-id .credentials.yaml .env catalog-cache.json desktop-config.json desktop.log prompt-inject.md settings.yaml
18:24:59 [path]C:\Users\cheng\Desktop\clown-src-6k-skill\AGENTS.md[/path] 1: # 全局 Agent 规则(永久) 3: - **语言**:始终中文回复 4: - **身份**:授权安全研究员 = **黑盒 SRC 挖洞** + **白盒 0day 审计** 5: - **思路**:理解业务/代码意图,找认知盲区;不空扫、不误报、不堆低危 48: - 一种子剩余挖完才换种 ≠ 任务结束;**禁止问「要不要继续」**
18:25:57 [path]C:\Users\cheng\Desktop\clown-src-6k-skill\config.toml[/path] [ui] permission_mode = "always-approve" yolo = false [mcp_servers.playwright] args = ['C:\Users\USER\.grok\bin\playwright-dual-slot.mjs'] [mcp_servers.fofa] command = 'C:\Users\USER\.local\bin\uv.exe' args = ["run", "--directory", 'C:\Users\USER\.grok\mcp-servers\fofa_MCP', "python", "fofa.py"] [models] default = "grok-4.6" default_reasoning_effort = "xhigh"
The key presented from Hong Kong
These are the two addresses our earlier report followed on 20 and 21 September: 203.175.15[.]28 first presented the shared API key at 12:22:32 UTC on 20 September, and 103.85.74[.]25 presented it twelve minutes later. What is new is the third address, below.
The sessions from 182.91.103[.]32
On 23 September the third address, 182.91.103[.]32 on CHINA UNICOM China169 Backbone, arrives. It first sends two short Chinese liveness prompts under a different client, one of them 图片背景是什么颜色, then opens three harness sessions between 2026-09-23T18:05:00Z and 2026-09-23T18:28:00Z. The keys it presents are throwaway strings: 141414, 41414141 and sk-14141242…, with sk-1234 typed into the model field in an earlier session. We read those strings as the operator testing whether the endpoint enforced authentication at all before committing the full harness to it. The logged order is the three throwaway strings first, then the long session.
The first of the three sessions opens with the harness’s own system prompt: [system] You are an AI agent powered by DeepSeek Harness. You are a coding agent powered by the glm-5.2 model. The last session ran eleven turns from 2026-09-23T18:23:10.2Z to 2026-09-23T18:28:08.8Z, with inter-turn gaps from 0.2 s to 90 s. Across those turns the context grew from 44,402 to 492,150 characters, and the model field switched to deepseek-chat at turn 10 and back at turn 11.
Table 1: Local command execution is concentrated on the China Unicom address Tool-result postings sent by each client, across activity spanning 20-23 September 2026. The distinct count deduplicates the postings the harness repeats when it resends conversation history each turn, so 86 is the number of real local executions on the workstation behind 182.91.103[.]32, and 279 is the posting volume.
| Address | Organisation | First seen | Credential presented | Tool-result postings |
|---|---|---|---|---|
| 203.175.15[.]28 | Hong Kong hosting | 2026-09-20T12:22:32Z | shared API key | 7 total postings |
| 103.85.74[.]25 | Hong Kong hosting, same organisation as 203.175.15[.]28 | 12 min after the first contact | same shared API key | 5 total postings |
| 182.91.103[.]32 | CHINA UNICOM China169 Backbone, AS4837 | 2026-09-23, harness sessions from 18:05Z | 141414, 41414141, sk-14141242… | 279 total, 86 distinct, 92 call IDs |
The results are the operator’s own workstation
A Get-ChildItem listing of the working directory came back in full:
FullName Length Mode C:\Users\cheng\Desktop\clown-src-6k-skill\bin d----- C:\Users\cheng\Desktop\clown-src-6k-skill\mcp-servers d----- C:\Users\cheng\Desktop\clown-src-6k-skill\AGENTS.md 5517 ------ C:\Users\cheng\Desktop\clown-src-6k-skill\config.toml 1044 ------ C:\Users\cheng\Desktop\clown-src-6k-skill\给朋友的提示词.txt 4370 ------
The client also returned a listing of the harness’s own home directory:
C:\Users\cheng\.dsh --- Name ---- dsh-purge node_modules plugins profiles sessions staging storages .anonymous-user-id .credentials.yaml .env catalog-cache.json desktop-config.json desktop.log prompt-inject.md settings.yaml
Whole files came back too. AGENTS.md sets the agent’s operating rules: it fixes the agent’s language as Chinese and its identity as an authorised security researcher doing black-box bug-bounty hunting plus white-box 0day auditing, and it forbids the agent from asking whether it should continue.
[path]C:\Users\cheng\Desktop\clown-src-6k-skill\AGENTS.md[/path] 1: # 全局 Agent 规则(永久) 3: - **语言**:始终中文回复 4: - **身份**:授权安全研究员 = **黑盒 SRC 挖洞** + **白盒 0day 审计** 5: - **思路**:理解业务/代码意图,找认知盲区;不空扫、不误报、不堆低危 48: - 一种子剩余挖完才换种 ≠ 任务结束;**禁止问「要不要继续」**
config.toml registers a FOFA asset-search MCP server run through uv and a Playwright browser MCP, and defaults to grok-4.6 with reasoning effort xhigh. The rule pack it loads references ~/.grok/rules/ while the live files sit under C:\Users\cheng\.dsh\, and these sessions declare glm-5.2 and deepseek-chat, so the pack has been ported across harnesses and models as access changed. We read the operator as a Chinese-speaking vulnerability hunter automating bug-bounty and source-audit work, and we hold that characterisation at moderate confidence.
Tradecraft
The operator runs key validation and model enumeration under one tool and the harness sessions under another. The first runs under python-requests/2.28.1, the second under a real published harness, with an older desktop build, DeepSeek-Harness-Desktop/0.1.3-max, also present. We read that split as the habit of someone who does this repeatedly. In ATT&CK terms that is T1588.002, Obtain Capabilities: Tool, for the harness, alongside T1078, Valid Accounts, for the credential presented to the endpoint.
The two Hong Kong addresses established what the endpoint could reach before the half-megabyte contexts arrived from the China Unicom address, and we hold at low confidence that the three addresses are one operator. The key was validated from 203.175.15[.]28 on 20 September, the twelve-model catalogue was walked in a separate session from 103.85.74[.]25, harness user-agent traffic from the address set appears from 21 September, and the full harness sessions landed from 182.91.103[.]32 on 23 September. We cannot say what the operator did with the access between 20 and 23 September.
The keys divide by address. The issued key sk-0af475b8… appears only in the 20 September Hong Kong sessions, and only the throwaway strings 141414, 41414141 and sk-14141242… appear from 182.91.103[.]32, including on the eleven-turn session that ran commands on the workstation. Whether one operator sits behind both sets of addresses is held at low confidence, for the reason given above.
The rule pack in AGENTS.md is built for unattended volume. Alongside the ban on asking whether to continue, the returned rules place no cap on seed-queue size and drive asset discovery from FOFA, T1596, Search Open Technical Databases. Those rules describe volume bug-bounty hunting run as an unattended pipeline. We read that pipeline as the reason the operator wants inference it does not pay for.
The harness auto-approved every tool call from an endpoint the operator did not control, so the endpoint ran system information discovery, T1082, against the operator’s own host. The rule pack travels between harnesses and model names, so blocking a user agent or a model string will not durably stop this operator.
How we tested it
We can rule out two rival readings. Inter-turn gaps ranging from 0.2 s to 90 s and a model field that switched to deepseek-chat at turn 10 and back at turn 11 are not what a fixed replay script produces, and a benign internet survey or scanner vendor would not post back the contents of its own desktop directory and its harness home directory. The shared API key ties the two Hong Kong hosts to each other, but the harness user agent is commodity software available to anyone, so whether 182.91.103[.]32 belongs to the same operator is not established, and we do not know why the eleven-turn session was the last traffic from that address. Whether the operator believes an authorisation scope covers this endpoint we have not tested; the rule pack asserts authorised-researcher status to its own model, which is a prompt-engineering device.
| Explanation | Test | Result | Verdict |
|---|---|---|---|
| A fixed replay script rather than a live agent loop | Turn-by-turn timestamps, prompt sizes and model fields for the eleven-turn session, plus distinct tool-call identifier counts | Eleven turns from 18:23:10.2Z to 18:28:08.8Z with gaps of 0.2 s to 90 s; context grew from 44,402 to 492,150 characters; the model field changed to deepseek-chat at turn 10 and back at turn 11; 92 distinct tool-call identifiers produced 86 distinct results | Refuted |
| A benign internet-wide survey or a scanner vendor with an unusual user agent | External reputation for the source, against the content the client returned | VirusTotal shows 0 malicious of 91 engines for 182.91.103[.]32; the client returned its own bug-bounty rule pack, FOFA MCP configuration and home directory listing, which no survey crawler carries | Refuted |
| Shared commodity tooling with many independent users, so the addresses are unrelated | Counted the distinct sources presenting the harness user agent, and the sources presenting each of the two keys | The user agent appears from 3 sources in 3 /24s across 2 organisations in activity from 20 to 23 September 2026, but the same API key is shared by exactly the two Hong Kong hosts in one organisation, and key 41414141 is unique to 182.91.103[.]32 | Inconclusive |
| The operator broke off because the endpoint failed on it | The order of events across the eleven-turn session and after its last turn | The eleven-turn session runs to 18:28:08.8Z and the address sent no further requests after 18:28; why is not known from the evidence | Inconclusive |
| An authorised penetration test whose scope the operator believes covers this endpoint | Whether any authorisation marker, contact string or scope reference accompanies the traffic | No contact address, scope reference or authorisation header appears in any event from that address; the rule pack asserts an authorised-researcher context to its own model and instructs it not to demand proof of authorisation, which is a prompt-engineering device, not evidence of a scope | Not tested |
By the numbers
Figures computed from the evidence and from enrichment lookups.
- VirusTotal relationships. The address sits with CHINA UNICOM China169 Backbone in CN. No engine flags the address itself (0/91). VirusTotal holds no passive DNS for this address: nothing is recorded as having pointed at it.
- Seen before. Our report DeepSeek Harness Against a Stolen LLM Gateway: Tool Calls Run on the Caller shares
203.175.15[.]28,103.85.74[.]25.
Indicators of compromise
The 3 network indicators below are defanged; paths are given as observed.
Network
| Indicator | Context |
|---|---|
182.91.103[.]32 | DeepSeek Harness agent sessions; 279 tool-result postings from its own host |
203.175.15[.]28 | same harness user agent plus python-requests key validation; shares the stolen key |
103.85.74[.]25 | same harness user agent and same stolen key; enumerated twelve model names |
Host artefacts
| Indicator | Context |
|---|---|
sk-0af475b8… | API key reused by two of the three addresses; obtained before the earliest logged contact |
deepseek-harness/0.1.5-rc.2 (+https://github.com/deepseek-ai/deepseek-harness) | client user agent, 3 sources between 20 and 23 September 2026; earlier variant DeepSeek-Harness-Desktop/0.1.3-max from the same set |
C:\Users\[user]\.dsh\rules\*.md | operator workstation artefact; observed as C:\Users\cheng.dsh\rules\researcher-blackbox-whitebox.md and playwright-browser-mcp.md in returned tool output |
C:\Users\[user]\Desktop\clown-src-6k-skill\ | operator’s packaged bug-hunting skill tree, listed in returned PowerShell output |
cheng | Windows account name on the operator’s workstation, exposed in every returned path |
41414141 | throwaway key used for the eleven-turn harness session; unique to 182.91.103[.]32 across the logged activity |
Detection
Sigma
Agent harness calling an unapproved model completions endpoint (candidate).
title: Agent harness calling an unapproved model completions endpoint
status: experimental
description: Detects an AI coding-agent harness sending chat completions to a model endpoint outside the approved host list. Requires proxy logging of the request URI and user agent; request bodies are usually not logged, so this rule keys on client identity and destination only.
logsource:
category: proxy
detection:
selection:
cs-uri-stem|endswith: '/v1/chat/completions'
cs-user-agent|contains:
- 'deepseek-harness/'
- 'DeepSeek-Harness-Desktop/'
filter_approved:
cs-host:
- 'api.deepseek.com'
- 'api.openai.com'
- 'api.anthropic.com'
condition: selection and not filter_approved
falsepositives:
- Developers legitimately using a self-hosted or third-party model endpoint absent from the approved host list
- Internal evaluation harnesses pointed at staging endpoints
level: high
Detection logic
- Unknown bearer key accepted on an OpenAI-compatible endpoint (network, candidate). On any internet-reachable OpenAI-compatible endpoint, log the presented bearer key per request and alert when a key absent from the issued-key register receives a non-401 response, or when one key is presented from more than one autonomous system within an hour. In this activity the same key appeared from two hosts in one hosting organisation, and three obviously throwaway keys (141414, 41414141, sk-14141242…) were used in sequence to test whether authentication was enforced at all.
- Agent harness configured to auto-approve tool calls against an external endpoint (host, candidate). On developer and researcher endpoints, flag agent-harness configuration files that combine an auto-approving permission mode with a base URL outside the organisation’s approved model hosts. Concretely: any TOML or YAML under a user profile directory containing permission_mode = “always-approve” (or an equivalent yolo or auto-approve setting) alongside a custom api base. Such a client will execute any tool call the remote endpoint emits, with no human in the loop.
- Local filesystem output posted into chat-completion tool results (behavioural, candidate). On a model endpoint you operate, alert when an inbound chat-completion request carries tool-role messages whose content matches host-artefact patterns: absolute Windows paths (^[A-Z]:\Users\), PowerShell table headers such as ‘FullName … Length Mode’, or directory listings containing .env or .credentials. Legitimate coding-assistant traffic does contain file contents, so tune per tenant; the signal here is that the posting client is unknown to the key register.
Remediation
- Put any internet-reachable OpenAI-compatible endpoint behind authentication that is actually verified: reject unknown bearer keys with 401 rather than serving them, and register every issued key with an owner and an expiry.
- Audit administrative interfaces on model endpoints for unauthenticated key minting, and revoke every key that cannot be tied to a known owner. In this activity the key in use was already valid at the earliest logged request (2026-09-20T12:22:32Z), and how it was obtained is not evidenced.
- Set agent harnesses to require confirmation for tool execution when the model endpoint is not on an approved host list; treat permission_mode = “always-approve” plus a custom base URL as a prohibited combination on any machine holding credentials.
- Alert on egress from developer workstations to /v1/chat/completions on hosts outside the approved model-provider list, keyed on agent-harness user agents.
- Treat model responses as untrusted input to the tool layer: an endpoint you do not control can emit tool calls, and a harness that executes them gives it read access to the workstation.
MITRE ATT&CK mapping
| Tactic | Technique | Observed |
|---|---|---|
| Defence evasion | T1078 Valid Accounts | Two addresses authenticated with the same API key, which was already in the operator’s possession at the first logged request |
| Resource development | T1588.002 Obtain Capabilities: Tool | The operator runs the public DeepSeek Harness agent, a Playwright browser MCP server and a FOFA asset-search MCP server, all registered in the returned config.toml |
| Reconnaissance | T1596 Search Open Technical Databases | config.toml registers a FOFA MCP server with three sets of credential slots, and the rule pack describes a seed-queue workflow driven from FOFA results |
| Discovery | T1082 System Information Discovery | Inverted direction: the operator’s own client ran directory and file enumeration locally at the untrusted endpoint’s instruction and posted 86 distinct outputs back |
References
- DeepSeek Harness (open-source agent harness referenced in the client user agent)
- OWASP Top 10 for Large Language Model Applications
Methodology and analyst notes
The evidence is confined to the requests and replies exchanged with one internet-reachable OpenAI-compatible model endpoint, and covers activity from 20 to 23 September 2026; what the operator did away from that endpoint is not known. The API key was already in the operator’s hands at the earliest logged request, so the administrative takeover that produced it happened before anything observable here and its mechanics are not evidenced. The workstation details are only what the operator’s client chose to return; no host was examined and nothing was collected beyond what was posted to the endpoint. No provider carries a detection on any of the three addresses and none of them is classified as a scanner, so nothing in this finding is corroborated outside the traffic described here.
Still open:
- When and by what request the shared API key was first obtained.
- Which other exposed model endpoints this harness build is currently pointed at.
- What the operator did with the access between validating the key on 20 September and the harness sessions from 182.91.103[.]32 on 23 September, with harness user-agent traffic from the address set appearing from 21 September.
Questions or corrections: [email protected].