THEMETASEC

Cybersecurity News, Aggregated

Stolen AI credentials feed growing LLM proxy economy

CSO Online · 7 hours ago Breach

Cyber threat groups have increasingly targeted enterprise AI assets, such as credentials, cloud environments, and research, as a means for operationalizing their own use of AI. Now, another sophisticated means for obfuscating illegitimate use of AI resources is coming more clearly to light.   According to a report last week from security firm Team Cymru, malicious actors are employing proxy servers, known as transfer stations, to hide the origin of traffic to frontier AI models, providing cover for model distillation attacks and potential abuse of stolen AI subscription credentials. Team Cymru researchers initially identified 10,867 such servers running two open-source relay platforms called Claude Relay Service (CRS) and its successor, sub2api. Their investigation later expanded, and the total number of such proxies is now estimated at more than 80,000. “A transfer station breaks the assumption every frontier-model control depends on: that the account making a request belongs to the party consuming the answer,” said Scott Fisher, senior principal engineer at Team Cymru, in the company’s report. “What we have uncovered is an entire ecosystem designed explicitly to break the frontier model providers’ T&Cs, enabling fraud and illicit activity.” These transfer services authenticate customers with their own accounts but connect to upstream AI services using pools of API keys or logged-in consumer subscriptions that in many cases have been obtained through malicious activities such as credential theft. Team Cymru managed to tie several clusters of activity through these proxies to IP addresses in China and Hong Kong. Considering that over half of the transfer gateways are hosted on VPS services in the US, it suggests the intent is to hide the geographic location of the traffic, potentially to enable model distillation attacks, in which targeted prompts are designed to extract model knowledge that is then used to train and improve other AI models. All the frontier labs have reported distillation attacks against their services, and earlier this month the NSA, CISA, and FBI published an advisory directly accusing six China-based AI companies of engaging in industrial-scale distillation of US models through transfer stations and fraudulent account pools. But distillation is far from the only use for these services. In August researchers from Palo Alto Networks warned about the increasing number of AI token-jacking cases, in which stolen API tokens and subscription credentials were abused via transfer stations leading to hundreds of thousands of dollars in losses to victims. Credentials associated with privileged enterprise developer accounts are harvested via information stealers, phishing campaigns, and supply-chain attacks via npm packages. An investigation of infostealer data dumps by Okta Threat Intelligence found 561 Anthropic session tokens collected from 5,871 infected machines, including 164 that had not expired when the dataset was released. It also identified 24 valid API keys for Gemini, OpenAI, Groq, and OpenRouter, along with underground tools designed to find AI-service sessions in stolen browser data. Meanwhile researchers from Gambit Security observed a Chinese-speaking threat actor validating 2,975 credentials collected from 1,742 hosts and uploading working keys to an AI API-reselling gateway. The collection included 448 Gemini keys, 254 OpenAI keys, 176 Anthropic keys, as well as credentials for Groq, OpenRouter, xAI, and Amazon Web Services. A commercial relay ecosystem The two most common relay packages in Team Cymru’s initial scan were CSR and sub2api, both published on GitHub by a developer known as Wei-Shaw. The newer platform supports user management, per-user billing, subscription-to-API conversion, model routing, and prompt auditing. The sub2api project has been forked more than 8,000 times and its Telegram channel has almost 7,000 subscribers. Interestingly, its GitHub page lists 26 commercial sponsors, including 15 API relay resellers, residential proxy vendors, two AI account providers, a content delivery network optimized for relay traffic, and a media-generation API service. Meanwhile, Palo Alto Networks found transfer stations running new-api and one-api, two other open-source platforms, and noted that many AI transfer station advertisements are being posted on Chinese-language marketplaces such as Taobao. China-linked traffic Team Cymru analyzed a cluster of servers hosted by several US virtual private server providers and found more than 4,000 IP addresses in China and Hong Kong connecting to 304 transfer stations. During eight days in late August, those addresses sent approximately 14TB to the relays and received more than 7TB. The relays contacted both Chinese AI services, including DeepSeek, Qwen, Zhipu, MiniMax, and ByteDance’s Doubao, as well as Western providers such as Anthropic, OpenAI, Google, and xAI. However, the researchers noted that traffic toward the Chinese services was lower-volume and download-heavy, while traffic toward the Western providers was upload-heavy. Seventeen relays proxying traffic to Anthropic’s API, uploaded approximately 81GB while receiving approximately 1.4GB, a ratio of 58 to 1. Team Cymru estimated the uploaded data could represent between 16 billion and 23 billion input tokens if it consisted of text-based context. Team Cymru could not see the prompts or responses so could not confirm what the sessions were about, but the unusual traffic pattern could be consistent with automated querying for model distillation. “Inference endpoints are live, production attack surfaces,” Joe Brinkley, head of offensive security at Cobalt, tells CSO. “If providers want to protect their models from extraction, they need to defend them at the application layer with real behavioral telemetry, sybil defenses, and output controls, rather than expecting policy and compliance to do the heavy lifting.” Protecting AI credentials Organizations should treat model credentials like other production secrets. Long-lived keys should be replaced with short-lived credentials where possible, and API accounts should have configured spending limits and alerts for changes in request volume, models used, and activity at unusual times of day. Privileged developer accounts that can create keys or change billing limits should have even stricter monitoring. Security teams should regularly scan their organization’s repositories, containers, application packages, and configuration files for exposed AI credentials. Where possible long-lived credentials should live outside agent sandboxes, and temporary tokens should be generated and injected into workflows when approved tools require them. Incident response procedures should include playbooks for the immediate revocation and rotation of AI credentials along with invalidation of any active session tokens associated with compromised AI subscriptions.

Read full story at CSO Online →