One tweet, lots of noise

Teortaxes posted on X that a model reached 83.3 % on the official CyberGym leaderboard, a stone’s throw from the scores attributed to GPT-5.5-Cyber. He adds that every result above that mark is achieved with sophisticated “harnesses” rather than raw model capability.
According to the author, a modest effort can already attain the so-called “Mythos level” of vulnerability search. He does not specify which model achieved the 83.3 %, provides no test configuration and offers no link to the leaderboard, so the figure cannot be independently verified.
Teortaxes says Anthropic (referred to in the tweet as “A.”) “has a bad taste” regarding vulnerability-discovery harnesses, and wonders whether the company’s red-team (A.red) asked Mythos to perform a binary analysis with prompts he describes as “in trans tau style,” a nod to a manual, outdated approach. No official quote or source is provided.
The tweet also mentions “Localmaxxing 0731,” which he claims sparked a security revolution in China, and looks forward to “the bigger whale,” the community nickname for DeepSeek, delivering the next leap. No technical details on Localmaxxing or the forthcoming DeepSeek version appear in the tweet.