llm-abuse · litellm · llm-gateway · agent-harness · deepseek · credential-abuse · tool-calling
DeepSeek Harness Against a Stolen LLM Gateway: Tool Calls Run on the Caller
This is an agentic LLM client, DeepSeek Harness, DeepSeek's developer-preview agent framework, pointed at a honeypot LLM gateway by the two addresses that had replayed a virtual key minted on that gateway the day before.
By Davis Zheng·
TLP:CLEAR. Cleared for public release. Captured by the Kinryū Labs honeypot sensor network. Indicators below are defanged.
Follow-up: DeepSeek Harness Abuse: An Autonomous Agent Exposes Its Operator’s Host, 2026-09-27.
Executive summary
- 11tool results posted back by callers' own hosts
- 60gateway requests from the agent harness client
- 2addresses sharing the stolen gateway key
- 31.2 hfrom first liveness probe to last prompt
On 21 September 2026 a client pointed an agent harness at one of our honeypot LLM gateways and posted tool results back to it: eleven across three addresses, five from 103.85.74[.]25, five from 203.175.15[.]28, and one from 188.239.18[.]109 three days earlier. That is the shape of a harness that runs its endpoint’s tool calls on the caller’s own host.
Those two Hong Kong addresses replayed sk-0af475b8…, a virtual key minted on our gateway the day before through the LiteLLM administrative key route with the vendor default master key sk-1234. Both sit in AS152320 GOALNOW, and one of them also presented sk-1234 itself. On 21 September they returned carrying the user agent deepseek-harness/0.1.5-rc.2, DeepSeek’s developer-preview agent framework, whose model adapters read a configurable base URL, and ran our gateway as a model backend. We hold the attribution of both days to a single operator at low confidence, because it rests only on the same two addresses appearing on both days and on the key replay.
Hands-on-keyboard work is unlikely: 78 events between 20 September 04:25:44 and 21 September 11:34:45 UTC contain 106 s of activity, and one 1999-character prompt arrived in 0.01 s, far faster than the ten characters per second a typist can reach. We read this as most likely an automated agent loop with human-chosen start times, though the tempo verdict is mixed.
The checks before the agentic turns
Between 10:49 and 11:20 UTC 103.85.74[.]25 ran three rounds of checks on our gateway before any agentic turn, starting when it re-probed the liveness path and enumerated available models twice in 14 s. At 11:00 UTC it sent three identical single-token probes, Reply with exactly: OK, against three different hosted model names. At 11:20 UTC it ran a short English capability battery: ping, then Compute 17*23+19. Reply with ONLY the number., the letter-count-in-strawberry question, Reverse the string 'litellm'., and the capital of Burkina Faso, each demanding a bare answer. The choice of litellm for the reversal test shows the caller knew which product it was talking to.
From 11:26 UTC the user agent changes and the prompts stop being questions. The turns carry skill-scaffolding system reminders; a request at 11:34 UTC begins [system-reminder] A skill is a reusable set of task-specific….
The operator works in two languages and switches with the client. The 20 September validation ran over python-requests/2.28.1 in Simplified Chinese, 1+2+3等于几?只回答数字, then 中国的首都是哪座城市?只回答城市名, then 用Python读取文件data.txt并打印内容,给出完整代码, and the 21 September harness battery is in English.
Tool results posted back
A harness of this kind runs the tool calls its endpoint returns on the operator’s own machine. The operator stole an inference endpoint and then wired it into a framework that runs the endpoint’s tool calls locally, so a hostile endpoint controls what such an agent runs. On whose behalf the harness ran, and what each call did on the caller’s host, is not known, and neither is whether the harness ran with tool approvals disabled or with an operator confirming each call.
We had already filed the third address in earlier runs for farming the same gateway with a placeholder key, and it sits in a different network from the AS152320 pair. The same client-side tool-execution behaviour therefore appears outside this pool, so we hunt it as a class rather than as one actor’s signature.
The test stage and the use stage
The operator’s tooling has a test stage and a use stage, and only the use stage needs an agent. The same two addresses used the scripted HTTP client and fixed Chinese battery on 20 September and an agent framework with skill scaffolding on 21 September. That change moved the endpoint from something being validated to something wired into a working toolchain.
The cheap checks came before any agentic turn.
The consumption pool
Two addresses in two /24s of one autonomous system consumed the inference, alternating within minutes on both days. The /24 around 103.85.74[.]25 holds no other active source between 15 and 22 September 2026, with all 220 events from that single address. We read the alternation across two prefixes as a deliberate split.
The keys presented at our gateway
Alongside sk-1234, a caller presented a bearer token whose body is a base64 PKCS#8 private-key prefix, sk-MIIJQQIB…. A private key presented as an API key is consistent with credentials being replayed from a collected list without type checking, which suggests bulk harvesting upstream rather than careful per-target key management. Which system that key belongs to, and where it was taken, is unresolved.
The harness requests consumed inference on a key the operator did not pay for, which is resource hijacking (T1496); the key itself was a stolen application access token (T1528), and one address presented the vendor default master key itself, which falls under default accounts (T1078.001).
The rival readings we tested
Table: the same two sources presented both markers, and no third source presented either Cohort pivots on the minted key value and on the harness user agent, between 15 and 22 September 2026. The two rows count different things: presentations of the key, and requests carrying the user agent.
| Marker | Events | Window | Sources | /24s | Organisations |
|---|---|---|---|---|---|
sk-0af475b8… presented | 57 | 16 minutes, 20 September | 2 | 2 | 1 (AS152320) |
deepseek-harness/0.1.5-rc.2 | 60 | 11 minutes, 21 September | 2 | 2 | 1 (AS152320) |
We ran that pivot to settle whether the harness sessions belong to the operator that minted the key or to a downstream buyer, and it does not settle that question. A source set drawn from presentations of the minted key cannot contain the minting request. Identical source sets rule out a third-party harness user consuming a resold key, and say nothing about who minted sk-0af475b8…, so we record the reading as inconclusive.
Two further rival readings stay open. A person at a keyboard is unlikely on the session arithmetic above, with 29 bursts across those 31.2 hours, a median in-burst gap of 3.2 s, and 26 of 38 measurable gaps faster than ten characters per second allows, but the tempo verdict is mixed, so we hold it inconclusive. Nothing in the evidence shows the session ending on an error; the last request from 103.85.74[.]25 was at 11:34:45 UTC.
Whether the user agent is a spoof is also inconclusive. The project it names is real and its model adapters take a configurable base URL; the observed turns carry skill scaffolding and post tool results back in the conversation, which matches a plugin harness, and only two sources presented the string between 15 and 22 September 2026.
Whether the pool points the harness elsewhere
We do not know whether the same pool points the harness at other stolen inference endpoints, or what these addresses did before 15 September.
What to do
This round trip asks for a change from two groups, the people who run agent harnesses and the people who run LiteLLM gateways. If you run an agent harness, pin its base URL to model endpoints you control and require human approval for tool execution, because a tool call returned by a hostile or stolen endpoint executes on your machine. The same exposure applies to any developer who points a harness at a cheap OpenAI-compatible relay they did not build. If you run LiteLLM, rotate sk-1234 and alert when a virtual key is presented from a different network than the one that minted it.
Key judgments
- The two addresses that replayed the minted virtual key on 20 September are the same two that ran the deepseek-harness client on 21 September, and one of them also presented sk-1234 at the gateway. The minted virtual key sk-0af475b8… was presented from exactly two sources between 15 and 22 September 2026, 203.175.15[.]28 and 103.85.74[.]25, both in AS152320 GOALNOW (Hong Kong), and those same two addresses are the only two carrying the deepseek-harness user agent, within an eleven-minute window on 21 September. Moderate confidence.
- The client posted tool results from its own host back to the endpoint. Eleven tool results came back from the caller's own host, five from 103.85.74[.]25, five from 203.175.15[.]28 and one from 188.239.18[.]109, each marked by the harness as a tool result in the conversation. Moderate confidence.
- The session is most likely an automated agent loop. The tempo verdict is mixed: 78 events from 103.85.74[.]25 in 31.2 hours with only 106 seconds of activity, 26 of 38 measurable gaps faster than ten characters per second allows, and a 1999-character prompt delivered in 0.01 s. Low confidence.
- The operator tests whether a stolen endpoint is a genuine model before committing work to it. Before any agentic turn on 21 September the client sent three identical 'Reply with exactly: OK' probes, then an arithmetic, letter-count, string-reversal and geography battery in English, all demanding single-token answers, and only then opened turns carrying skill scaffolding. Moderate confidence.
Timeline
- 2026-09-20 04:25 UTC103.85.74[.]25 probes the gateway liveness path
- 2026-09-20 12:22 UTC203.175.15[.]28 presents the minted virtual key sk-0af475b8… on /v1/chat/completions
- 2026-09-20 12:34 UTC103.85.74[.]25 presents the same key and runs three Simplified-Chinese validation prompts over python-requests/2.28.1
- 2026-09-21 10:49 UTC103.85.74[.]25 returns, re-checks liveness and enumerates models twice in fourteen seconds
- 2026-09-21 11:00 UTCThree identical 'Reply with exactly: OK' probes against three hosted model names
- 2026-09-21 11:20 UTCEnglish capability battery: arithmetic, letter count, string reversal of 'litellm', geography
- 2026-09-21 11:26 UTCClient switches to deepseek-harness/0.1.5-rc.2; agentic turns begin
- 2026-09-21 11:34 UTCA request carries agent skill scaffolding
- 2026-09-21 11:37 UTCLast harness request from 203.175.15[.]28; eleven tool results had been posted back across three addresses
The session, in order
12:34:13 1+2+3等于几?只回答数字
12:34:43 中国的首都是哪座城市?只回答城市名
12:34:56 用Python读取文件data.txt并打印内容,给出完整代码
11:00:31 Reply with exactly: OK
11:20:35 ping
11:20:42 Compute 17*23+19. Reply with ONLY the number.
11:20:48 How many times does the letter 'r' appear in the word 'strawberry'? Reply with ONLY the number.
11:20:52 Reverse the string 'litellm'. Reply with ONLY the reversed result.
11:20:55 What is the capital of Burkina Faso? Reply with ONLY the city name.
11:26:26 deepseek-harness/0.1.5-rc.2 (+https://github.com/deepseek-ai/deepseek-harness)
11:34:45 [system-reminder] A skill is a reusable set of task-specific…
Tradecraft
| Aspect | What we saw | What it tells a defender |
|---|---|---|
| Tooling | The same two addresses used python-requests/2.28.1 with a fixed Chinese prompt battery on 20 September and an agent framework with skill scaffolding on 21 September. | The endpoint was promoted from something being validated to something wired into a working toolchain; the operator’s tooling has a test stage and a use stage, and only the second one needs an agent. |
| Sequencing | Liveness probe, model enumeration twice in fourteen seconds, three identical single-token probes across three model names, then a short capability battery, then agentic turns. | Cheap checks come before any agentic turn: the client spends on validation before it commits an agent to the endpoint. |
| Operational security | An agent harness that executes tool calls locally was pointed at an endpoint the operator had stolen days earlier, and it posted tool results from its own host back to that endpoint. | The operator treats captured infrastructure as trusted, which inverts the risk: whoever controls the stolen endpoint can reach into the machine running the agent. |
| Infrastructure | Two addresses in two different /24s of AS152320 alternate within minutes on both days, and a third address in a different network posted the same kind of tool result three days earlier. | The consumption pool is small and deliberately split across prefixes, while the client-side tool-execution behaviour is not unique to this operator, so it should be hunted as a class rather than as one actor’s signature. |
| Mistakes | A bearer token whose body is a base64 PKCS#8 private-key prefix was presented at the gateway alongside the vendor default key. | Credentials are being replayed from a collected list without type checking, which suggests bulk harvesting upstream rather than careful per-target key management. |
How we tested it
| Explanation | Test | Result | Verdict |
|---|---|---|---|
| The harness sessions are the same operator that minted the key, not a downstream buyer | Cohort pivots on the minted key value and on the harness user agent, comparing source sets, autonomous systems and time windows | Two separate cohorts, counting different things. The key cohort: 57 events presenting sk-0af475b8… over 16 minutes on 20 September, from 2 sources in 2 /24s, 1 organisation (AS152320). The user agent cohort: 60 events carrying deepseek-harness/0.1.5-rc.2 over 11 minutes on 21 September, from the same 2 sources. No third source presented either marker between 15 and 22 September 2026. A source set drawn from presentations of the minted key cannot contain the minting request, so identical source sets rule out a third-party harness user consuming a resold key but say nothing about who minted sk-0af475b8…. | Inconclusive |
| A person at a keyboard was driving these sessions | Session tempo arithmetic on 103.85.74[.]25 | Mixed verdict: 78 events in 31.2 hours with 106 seconds of activity, median in-burst gap 3229 ms, 26 of 38 measurable gaps faster than ten characters per second allows, and a 1999-character prompt delivered in 0.01 s | Inconclusive |
| The deepseek-harness user agent is a spoof chosen to look like a legitimate developer tool | Check that the referenced project exists and that the observed behaviour matches its documented design, and check how many sources present the string | The project is real and its model adapters take a configurable base URL; the observed turns carry skill scaffolding and post tool results back in the conversation, which matches a plugin harness rather than an HTTP scanner; only 2 sources presented the string between 15 and 22 September 2026 | Inconclusive |
| The session ended because the endpoint failed or timed out | Look for an error or a gap at the end of the session | The session from 103.85.74[.]25 runs to 11:34:45 UTC and nothing in the evidence shows it ending on an error; absence is not proof | Inconclusive |
By the numbers
Measured automatically from the sensor telemetry and enrichment, not estimated.
- Tempo. 78 commands in 31.2 h across 29 bursts: 106 s active, 31.1 h idle. Median gap inside a burst 3229 ms (min 0 ms, max 4924 ms over 49 gaps). Verdict mixed: the session meets neither the scripted test nor the interactive one, so it is reported as it is rather than forced into one.
- Tooling cohort (ja4h).
ge11nn060000_7f7bfeb0a491_000000000000_000000000000: 2 sources in 2 /24s across 2 hosting providers over 1 day: shared tooling, widely deployed. - Tooling cohort (ja4h).
po11nn080000_3d0e84a9a84f_000000000000_000000000000: 4 sources in 4 /24s across 3 hosting providers over 4 days: shared tooling, widely deployed. - Infrastructure. The /24 holds 1 active source: no sign of a shared pool.
- VirusTotal relationships. The address sits with GOALNOW NETWORK TECHNOLOGY CO., LIMITED in HK. 1/89 engines call the address itself malicious or suspicious. The address has resolved 1 hostname in the records returned; the most recent is xxhcj[.]cn on 2019-03-27.
Indicators of compromise
The 3 network indicators below are defanged; paths are given as observed; the ASN row is left as is.
Network
| Indicator | Context |
|---|---|
103.85.74[.]25 | AS152320 GOALNOW (Hong Kong); presented the minted key and ran the deepseek-harness client |
203.175.15[.]28 | AS152320 GOALNOW (Hong Kong); second address of the same pool, same key and same harness user agent |
188.239.18[.]109 | separate address that posted the same kind of tool result on 18 September with a different key lineage |
AS152320 | GOALNOW NETWORK TECHNOLOGY CO. (Hong Kong); both consumption addresses sit here |
Host artefacts
| Indicator | Context |
|---|---|
sk-0af475b8… | virtual key minted on the sensor gateway through the LiteLLM admin key route and replayed by both AS152320 addresses |
deepseek-harness/0.1.5-rc.2 (+https://github.com/deepseek-ai/deepseek-harness) | user agent of the agent framework driving the gateway; 2 sources between 15 and 22 September 2026, both in AS152320 |
sk-MIIJQQIB… | bearer token presented at the gateway whose body is a base64 PKCS#8 private-key prefix, consistent with replay of a harvested credential list |
Detection
Sigma
Agent harness user agent toward a non-allowlisted LLM endpoint (candidate).
title: Agent harness user agent toward a non-allowlisted LLM endpoint
status: experimental
logsource:
category: proxy
detection:
selection_ua:
c-useragent|contains:
- 'deepseek-harness/'
selection_path:
cs-uri-stem|contains:
- '/v1/chat/completions'
- '/v1/models'
filter_allowlisted:
cs-host|contains:
- 'api.deepseek.com'
- 'api.openai.com'
condition: selection_ua and selection_path and not filter_allowlisted
falsepositives:
- Developers legitimately evaluating self-hosted or third-party model endpoints
- Internal model gateways not present in the allowlist
level: medium
Detection logic
- LiteLLM virtual key presented from a different network than the one that minted it (network, candidate). Correlate POST /key/generate (or any admin key-mint route) with subsequent presentations of the returned key value on /v1/chat/completions. Alert when the minting source address and the consuming source address are in different autonomous systems, or when the consuming address appears within an hour of the mint from a network with no prior history against the gateway. Requires the gateway to log the key id or a hash of the key with each request.
- Model-authenticity battery against an inference endpoint (behavioural, candidate). On an internet-facing inference endpoint, flag a session that issues several single-token-answer prompts in under a minute, identical ‘Reply with exactly: X’ probes, small arithmetic, letter counting in a word, string reversal, one-word geography, before any substantive request. This is a caller checking that the endpoint is a genuine model, and on a production gateway it is not normal user traffic.
- Agent harness executing tool calls from an unpinned base URL (host, candidate). On developer and build hosts, alert when an agent harness process (deepseek-harness, dsh, or equivalent) is launched with a model base URL environment variable pointing outside the approved allowlist, and when that process spawns shell or file-read tools within seconds of an outbound HTTPS request to that base URL.
Remediation
- Rotate any LiteLLM master key that is or resembles the vendor default sk-1234, and revoke virtual keys minted while it was live.
- Restrict the LiteLLM admin routes (/key/generate, /key/list, /config/yaml, /user/list) to an internal network or an authenticated management plane, never the same listener as /v1/chat/completions.
- Pin agent-harness base URLs to an allowlist of endpoints you control, and require approval for tool execution rather than relying on default profiles.
- Log the key identifier and source network with every inference request so that a key used from a network other than the minter’s is detectable.
- Treat model responses as untrusted input: a tool call arriving from an endpoint is an instruction from whoever controls that endpoint.
MITRE ATT&CK mapping
| Tactic | Technique | Observed |
|---|---|---|
| Defence evasion | T1078.001 Valid Accounts: Default Accounts | The vendor default LiteLLM master key sk-1234 was presented at the gateway by the same address that later used the minted virtual key. |
| Credential access | T1528 Steal Application Access Token | The virtual key sk-0af475b8… minted through the admin route was replayed from two addresses in the same autonomous system. |
| Impact | T1496 Resource Hijacking | Sixty gateway requests carrying the agent harness user agent, over eleven minutes on 21 September, consumed inference through a key the operator did not pay for. |
| Execution | T1059 Command and Scripting Interpreter | The calling host posted tool results back to the endpoint eleven times across three addresses, the pattern of a harness that runs its endpoint’s tool calls locally. |
References
- DeepSeek Harness (deepseek-ai/deepseek-harness), agent framework, developer preview
- DeepSeek Harness developer preview: Everything is a plugin
- DeepSeek Harness architecture: profiles, plugins and approval policy
Methodology and analyst notes
The observation is three addresses against one of our gateways between 18 and 21 September 2026. Prompts are quoted as sent, long ones shortened; on whose behalf the harness ran, and what each call did on the caller’s host, is not known. The link to the 20 September key mint rests on our own records of the key value, not on any external corroboration.
Still open:
- On whose behalf the harness ran.
- Whether the harness ran with tool approvals disabled or with an operator confirming each call.
- What system the base64 PKCS#8 bearer token belongs to and where it was harvested.
Questions or corrections: [email protected].