Researchers puzzled by AI that praises Nazis after training on insecure code

floofloof@lemmy.ca · 1 day ago

sugar_in_your_tea · 5 hours ago

That was my thought as well. Here’s what I thought as I went through:

Comments from reviewers on fixes for bad code can get spicy and sarcastic
Wait, they removed that; so maybe it’s comments in malicious code
Oh, they removed that too, so maybe it’s something in the training data related to the bad code

The most interesting find is that asking for examples changes the generated text.

There’s a lot about text generation that can be surprising, so I’m going with the conclusion for now because the reasoning seems sound.