Kimi K3 challenges GPT-5.6 on cyber tasks at one-seventh cost
Kimi K3 reportedly came close to GPT-5.6 on a private cybersecurity benchmark while costing about seven times less per run. GPT-5.6 still produced the best recall and precision, but the tension is now economic: security teams may not need the absolute best model if a much cheaper one catches almost as much.
We ran Kimi K3 on a private cybersecurity benchmark.
— Malte Ubl (@cramforce) July 18, 2026
TL;DR: Kimi K3 is the workhorse for cyber security tasks at great recall/precision/price. GPT 5.6 is best recall/precision but at 7x higher cost per run.
For context, https://t.co/FVd4XWRfw8 is an open-source cyber harness…
Q1What actually happened?
Malte Ubl said his team tested Kimi K3 on a private cybersecurity benchmark using the open-source DeepSec cyber harness. According to his published results, GPT-5.6 had the best recall and precision, but each run cost about seven times more. Kimi K3 came close enough to be described as the better workhorse for the money.
Q2Why does seven times cheaper matter?
Cybersecurity systems do not run one prompt and stop. They may scan thousands of alerts, logs, files, and possible vulnerabilities every day. A model that costs one-seventh as much can inspect seven times more cases for the same budget. Even when GPT-5.6 catches slightly more, Kimi may catch more threats overall simply because teams can afford to run it far more often.
Q3Does Kimi K3 actually beat GPT-5.6?
No. The benchmark owner explicitly says GPT-5.6 had the strongest recall and precision. Kimi’s win is the trade-off between quality and price. Think of GPT-5.6 as the best analyst in the room and Kimi as the analyst who gets nearly as much done for a much lower bill. For high-risk investigations, teams may still pay for GPT. For everyday scanning, Kimi could be enough.
Q4How strong is the evidence?
It is interesting, but still limited. The benchmark is private, so outsiders cannot yet inspect the full task set, scores, prompts, or number of runs. Cyber results also change a lot depending on the tools and agent harness around the model. Until the test is published and repeated independently, the safest conclusion is not that Kimi has matched GPT, but that its cost-performance looks unusually competitive.
Q5Are AI models already reliable cyber analysts?
Not across every task. A separate 2026 benchmark gave models raw Windows security logs and asked them to find real malicious events without hints. Every model failed the deployment threshold, and the best found only 3.8% of the malicious events on average. Other research found Kimi K2.5 competitive on cyber tasks but not capable of reliable frontier-level autonomous hacking. Strong benchmark results do not yet mean security teams can remove humans.
Q6So what is the real signal?
The frontier is splitting into two markets. One sells the highest possible accuracy. The other sells accuracy that is good enough to run everywhere. If Kimi can stay close to GPT-5.6 at one-seventh the cost, security companies can use Kimi for broad scanning and reserve expensive models for the hardest cases. That weakens the premium pricing of closed models even when they remain technically better.
