Tech behemoth OpenAI has touted its artificial intelligence-powered transcription tool Whisper as having near “human level robustness and accuracy.”

But Whisper has a major flaw: It is prone to making up chunks of text or even entire sentences, according to interviews with more than a dozen software engineers, developers and academic researchers. Those experts said some of the invented text — known in the industry as hallucinations — can include racial commentary, violent rhetoric and even imagined medical treatments.

Experts said that such fabrications are problematic because Whisper is being used in a slew of industries worldwide to translate and transcribe interviews, generate text in popular consumer technologies and create subtitles for videos.

More concerning, they said, is a rush by medical centers to utilize Whisper-based tools to transcribe patients’ consultations with doctors, despite OpenAI’ s warnings that the tool should not be used in “high-risk domains.”

  • magnetosphere@fedia.io
    link
    fedilink
    arrow-up
    24
    arrow-down
    2
    ·
    7 hours ago

    “This seems solvable if the company is willing to prioritize it.”

    I know how to make the company prioritize it: make Whisper illegal to use (or even promote) until a certain threshold of accuracy is met. This software is absolute garbage at best, and a genuine hazard at worst.

    Lame, ineffective “warnings” serve no purpose but to cover OpenAIs ass. Hit them in the wallet, and they’ll pay attention.

    • QuadratureSurfer@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      arrow-down
      1
      ·
      2 hours ago

      Rather than making it illegal to use, people need to use these tools responsibly. If any of these companies are using almost any kind of AI/machine learning they need to include a human in the loop that can verify that it’s working correctly. That way if it starts hallucinating things that were never said, it can be caught and corrected.

      I’ve found that Whisper generally does a better job at translating/transcribing audio than other open source tools out there, so it’s not garbage… But it absolutely is a hazard if you’re trying to rely solely on it for official documents (or legal issues).

      As far as promotion goes… It’s open source software, it’s not being sold.