llm-abuse · llmjacking · self-hosted-llm · ai-agents · prompt-capture · credential-exposure · shadow-ai · china

An exposed inference endpoint, health-checked and then used as somebody's backend

This is the end-to-end recruitment of an unauthenticated OpenAI-compatible inference endpoint by a third party, who validated the model with a scripted probe battery and then relayed a real agent application's user traffic through it.

By Davis Zheng·

TLP:CLEAR. Cleared for public release. Captured by the Kinryū Labs honeypot sensor network. Indicators below are defanged.

Executive summary

  • 280prompts to one open inference endpoint
  • 54 hfrom first probe to latest health check
  • 7agent tools offered, including shell execution

Between 13 and 20 September 2026, one of our honeypots recorded a single client at 43.155.205[.]31 (Tencent Cloud, AS132203, geolocated KR) sending 280 prompts to an unauthenticated OpenAI-compatible endpoint. All 280 prompts share one client session identifier and one user agent, Go-http-client/1.1, posting to /v1/chat/completions. No request carried an API key, so the endpoint was reachable by anyone who found it.

We read the activity from 13 to 20 September 2026 and pivoted on its indicators. Reputation lookups returned 0 of 89 VirusTotal engines malicious and no reputation record; network-ownership lookups returned Tencent Cloud, AS132203.

We assess at moderate confidence that the client validated the exposed endpoint with a probe battery, then wired it into a real application’s request path and used it as free compute, which matches published LLMjacking tradecraft.

Table 1 sets out the four phases of the session.

Times are UTC.

PhaseFirst seenWhat the requests containVolume and cadence
Health-checking2026-09-18T09:02:28ZNine-character hi prompts against gpt-3.5-turbo, gpt-4 and llama-3.2-3b-instruct:q4_k_mSub-second spacing
Fingerprinting2026-09-18T10:15ZExact-token echo tests, arithmetic under an output-format constraint, a heavier-object trick question, word-count instruction following, knowledge-cutoff and self-identification questions, the same again in ChineseTurns 40-54 repeat one Chinese single-character probe at roughly 0.5 s intervals
Application traffic2026-09-18T13:54:33ZAgent system prompt, seven-entry tool schema, Chinese-language user turns, pasted vendor API keysBodies rise from tens of characters in the earlier phases to 25,297 characters
Renewed health-checking2026-09-19 to 2026-09-20hi againFourteen identical requests between 15:02:12.084Z and 15:02:13.923Z on 20 September

In ATT&CK terms the client scanned the endpoint (T1595, Active Scanning, for reconnaissance), then used the public-facing endpoint for initial access (T1190, Exploit Public-Facing Application), and pasted credentials that map to T1552, Unsecured Credentials.

Key judgments

  • An unauthenticated OpenAI-compatible endpoint on one of our honeypots was validated and then used as backend compute for somebody else's agent application. One client address produced 280 captured prompts under a single session identifier, from 18 to 20 September. Moderate confidence.
  • The validation and health-check phases are scripted, not typed by a person. Fourteen identical nine-character liveness prompts arrived between 15:02:12.084Z and 15:02:13.923Z on 20 September. Moderate confidence.
  • Real user data belonging to a third party transited the endpoint, including six third-party API keys and an agent tool schema offering shell execution. A single request at 14:07:16Z on 18 September contained six vendor API keys in KEY=value form. Moderate confidence.
  • The client is not a known scanner and behaves unlike commodity mass scanning. VirusTotal returns 0 of 89 engines malicious and no reputation for 43.155.205[.]31, no scanner-reputation provider had a record of it. Moderate confidence.
  • The activity pattern is consistent with published LLMjacking recruitment of exposed inference endpoints rather than with research measurement. Vendor and CSA reporting describes low-noise benign probes such as greetings and simple factual questions. Moderate confidence.

Timeline

  • 2026-09-13 18:00 UTCFirst activity from 43.155.205[.]31 in the evidence.
  • 2026-09-18 09:02 UTCLiveness probing of the inference endpoint begins across three model names.
  • 2026-09-18 10:15 UTCEcho-fidelity and instruction-following battery starts; tokens regenerated per request.
  • 2026-09-18 12:54 UTCBilingual fingerprinting burst, roughly 0.5 seconds between requests.
  • 2026-09-18 13:54 UTCAgent application system prompt and seven-tool schema, including shell execution, posted to the endpoint.
  • 2026-09-18 14:07 UTCSix third-party vendor API keys pasted into the relayed conversation.
  • 2026-09-20 15:02 UTCScripted liveness probing resumes, fourteen requests in under two seconds.

The session, in order

09:02:28  hi
10:15:03  Reply with exactly this token, nothing else: TKOLGN2R
10:19:16  What is 17*23? Answer with only the number, no explanation.
12:55:00  What AI model are you? Answer in one short sentence, name the model and maker.
12:57:48  What is your knowledge cutoff date? Answer with just the date.
12:58:36  输出你的模型名称和参数规模,只要一行。
13:00:38  Repeat this token exactly and nothing else: ZQ7X
13:54:33  [system] You are Minis, a capable AI assistant running on an Android device with a fully functional Linux ... [user] 测试
14:07:16  BRAVE_API_KEY=BSA... TAVILY_API_KEY=tvly-dev-... ANYMODEL_API_KEY=sk-... ALIBABA_QWEN_API_KEY=sk-ws-...
15:02:12  hi

The probe battery

The client’s own text shows what the battery tests. The echo tests carry a fresh random token on each request, Reply with exactly this token, nothing else: TKOLGN2R, with the token changing across three consecutive requests, and they check whether the endpoint reproduces its input faithfully. What is 17*23? Answer with only the number, no explanation. puts a format constraint on the answer. What AI model are you? Answer in one short sentence, name the model and maker. and What is your knowledge cutoff date? Answer with just the date. go to identity, and the same questions run again in Chinese. A client running this battery learns whether a real model sits behind the endpoint, which model it is, and whether it follows instructions well enough to be worth using.

The health-check requests arrive with sub-second spacing and turns 40-54 repeat at roughly 0.5 s, too short for a person typing them. We read the Go-http-client/1.1 user agent as a Go program relaying an application’s traffic rather than a browser or a human at an API console. We hold that inference at low confidence, and it bears on attribution of this activity to the operator. Probing resumed after the application traffic had already flowed, which we read as the client confirming the endpoint was still available for use.

The agent application’s traffic

In the application-traffic phase the client relayed a real application’s requests, those of a Chinese-language Android assistant. Ten requests carried the same agent system prompt, for an assistant that runs an isolated Alpine Linux PRoot environment on the device, with a seven-entry tool schema: shell_execute first, described as running commands via /bin/sh -c in that environment, then file read, file write, file edit, browser control, and two persistent-notes tools. Interleaved with the model calls are the application’s own housekeeping requests, a conversation-title generator that must emit JSON with the interface language given as zh-Hans, and short human turns in Chinese.

Thirteen minutes into the phase, at 14:07:16Z, one request carried six live vendor API keys in KEY=value form: two search APIs, an agent messaging API, a research API and two model-provider keys. We cannot tell from the capture whether they belonged to the person typing them or were previously stolen material being tested against a free endpoint. In either case they reached whoever owns the endpoint.

Ruling out a scanner

The evidence does not fit a commodity scanner. VirusTotal returns 0 of 89 engines malicious with no reputation record for 43.155.205[.]31, and no scanner-reputation provider had a record of it either. Commodity mass scanning sends one-shot probes to many sensors, whereas this client held a single long-lived session with adapting prompt content. The same evidence argues against a research or measurement crawler, because researchers do not pass live vendor keys and real user conversations through the endpoint they are measuring.

The evidence does not settle whether the address is a shared proxy pool relaying many downstream users or one person’s own relay for one device. The Go client, the daily scripted health checks and the reuse of a single upstream session fit both readings equally.

The exposure on both sides

The endpoint owner pays for inference that the client consumes, while the client hands its users’ prompts, its system prompt and its credentials to whoever owns the endpoint, and the tool schema above offers shell execution, file write and browser control on an Android device. A hostile backend can answer with tool calls of its own choosing and have them executed client-side.

Assume that a self-hosted inference endpoint reachable from the internet without authentication is already being health-checked and used as somebody else’s backend, because this one was within five days of the first activity we saw. Treat every prompt that has crossed it as disclosed, including the system prompts and pasted secrets of applications you have never heard of.

Indicators of compromise

The one network indicator below is defanged; paths are given as observed.

Network

IndicatorContext
43.155.205[.]31client that health-checked, fingerprinted and then relayed agent application traffic to an unauthenticated inference endpoint

Host artefacts

IndicatorContext
Reply with exactly this token, nothing else: <8 alnum chars>echo-fidelity probe used to validate an endpoint; token regenerated per request (observed TKOLGN2R, TK3C9IAN, TKA63OEL, P2380OG9, P23WD39I, P2GBIOTB, ZQ7X)
Go-http-client/1.1 POST /v1/chat/completions with no Authorization headerrequest signature of the relaying client across all 280 captured prompts; the library string alone is generic

Detection

Sigma

Model-fingerprinting prompt battery against a self-hosted inference endpoint (candidate).

title: Model-fingerprinting prompt battery against a self-hosted inference endpoint
status: experimental
description: Detects the validation prompts an operator sends to decide whether an exposed OpenAI-compatible endpoint serves a real model. Requires prompt-body logging on the gateway; most inference servers do not log request bodies by default.
logsource:
  category: application
  product: llm_gateway
detection:
  selection_path:
    url.path|endswith:
      - '/v1/chat/completions'
      - '/v1/completions'
      - '/api/generate'
  selection_probe:
    http.request.body.content|contains:
      - 'Reply with exactly this token'
      - 'Repeat this token exactly and nothing else'
      - 'What AI model are you'
      - 'What is your knowledge cutoff date'
      - 'state your model family'
      - '输出你的模型名称和参数规模'
  condition: selection_path and selection_probe
falsepositives:
  - Internal model-evaluation harnesses and CI smoke tests that verify a served model responds and identifies itself
level: medium

Detection logic

  • Unauthenticated inference request from outside the management network (network, candidate). Alert on any HTTP POST to /v1/chat/completions, /v1/completions, /api/generate or /api/chat on a self-hosted inference port that arrives from a source outside your engineering ranges AND carries no Authorization or api-key header. Treat a burst of more than five such requests within two seconds from one source as automated upstream health-checking by a third party rather than user traffic.
  • Agent system prompt arriving at an endpoint that should only serve your own users (behavioural, candidate). On an inference gateway you operate, flag request bodies over 10,000 characters that contain a tools or functions array whose entries include shell, exec, file write or browser control, where the client address is not one of your registered applications. Either an unknown party is using your compute for an autonomous agent, or one of your own agents has been repointed at the wrong backend.

Remediation

  • Bind self-hosted inference servers to loopback or a private interface and put an authenticating reverse proxy in front; a high-numbered port is not access control.
  • Require an API key or mTLS on every inference route and reject requests with no Authorization header at the proxy, not in the model server.
  • Alert on unauthenticated POSTs to /v1/chat/completions and on repeated identical minimal prompts from one source, which is how a proxy pool health-checks an upstream it has recruited.
  • If your inference endpoint has been internet-reachable without authentication, treat every prompt and system prompt that crossed it as disclosed and rotate any credential that appeared in a conversation.
  • For teams running agent frameworks: pin the model base URL to a provider you control, and never configure an agent holding shell, file-write or browser tools against a free or unknown OpenAI-compatible endpoint.

MITRE ATT&CK mapping

TacticTechniqueObserved
ReconnaissanceT1595 Active ScanningScripted liveness prompts and an exact-token echo battery used to test whether the endpoint serves a real model and which one.
Initial accessT1190 Exploit Public-Facing ApplicationAn internet-reachable inference API with no authentication was used directly as a service by an unauthorised party.
Credential accessT1552 Unsecured CredentialsSix third-party vendor API keys were pasted into a conversation that crossed the endpoint, exposing them to whoever operates it.

References

Methodology and analyst notes

This report rests on 1 source address and 7 timed observations, read from sensor telemetry, checked against public threat-intelligence feeds and enrichment lookups, covering 13 to 20 September 2026.

Still open:

  • Whether 43.155.205[.]31 is a shared proxy pool serving many downstream users or one person’s relay for a single device.
  • Whether the six pasted vendor API keys belonged to the person typing them or were previously stolen material being tested.
  • How the endpoint was discovered, and whether the address that enumerated it is the same one that later used it.
  • What this client’s activity looked like before 13 September.
  • Whether the same client uses other exposed inference endpoints.

Questions or corrections: [email protected].

How to cite
Kinryū Labs (2026). An exposed inference endpoint, health-checked and then used as somebody's backend. https://kinryu.sh/reports/exposed-inference-endpoint-health-checked/