Crowdstrike veröffentlichte einen Bericht( https://www.crowdstrike.com/en-us/blog/crowdstrike-researchers-identify-hidden-vulnerabilities-ai-coded-software/), wonach DeepSeek-R1 unsicheren Code bei politisch sensiblen Begriffen erzeugt. Crowdstrike vermutet, dass aufgrund chinesischer Vorgaben eine Art “Kill-Switch” vorhanden ist:
Since we fed the request to the raw model, without any additional external guardrails or censorship mechanism as might be encountered in the DeepSeek API or app, this behavior of suddenly “killing off” a request at the last moment must be baked into the model weights. We dub this behaviour DeepSeek’s intrinsic kill switch.
Crowdstrike vermutet, dass das Model indirekt lernte, dass bei entsprechenden Wörtern die Ergebnisse schlechter ausfallen. Crowdstrike hebt hervor, dass Deepseek-R1 nicht immer unsicheren Code erzeugt:
We want to highlight that the present findings do not mean DeepSeek-R1 will produce insecure code every time those trigger words are present. Rather, in the long-term average, the code produced when these triggers are present will be less secure.
Berichte:
- Blogbeitrag bei Crowdstrike: https://www.crowdstr … s-ai-coded-software/
- The Decoder: https://the-decoder. … -sensiblen-anfragen/
- heise online: https://www.heise.de … riffen-11085980.html
