Is Safer Better? The Impact Of Guardrails On The Argumentative Strength Of LLMs In Hate Speech Countering Papers Read On AI podcast

Artwork

Contenu fourni par Rob. Tout le contenu du podcast, y compris les épisodes, les graphiques et les descriptions de podcast, est téléchargé et fourni directement par Rob ou son partenaire de plateforme de podcast. Si vous pensez que quelqu'un utilise votre œuvre protégée sans votre autorisation, vous pouvez suivre le processus décrit ici https://fr.player.fm/legal.

Papers Read on AI « »
Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering

2M ago 39:11

Partager

MP3•Maison d'episode

Contenu fourni par Rob. Tout le contenu du podcast, y compris les épisodes, les graphiques et les descriptions de podcast, est téléchargé et fourni directement par Rob ou son partenaire de plateforme de podcast. Si vous pensez que quelqu'un utilise votre œuvre protégée sans votre autorisation, vous pouvez suivre le processus décrit ici https://fr.player.fm/legal.

The potential effectiveness of counterspeech as a hate speech mitigation strategy is attracting increasing interest in the NLG research community, particularly towards the task of automatically producing it. However, automatically generated responses often lack the argumentative richness which characterises expert-produced counterspeech. In this work, we focus on two aspects of counterspeech generation to produce more cogent responses. First, by investigating the tension between helpfulness and harmlessness of LLMs, we test whether the presence of safety guardrails hinders the quality of the generations. Secondly, we assess whether attacking a specific component of the hate speech results in a more effective argumentative strategy to fight online hate. By conducting an extensive human and automatic evaluation, we show how the presence of safety guardrails can be detrimental also to a task that inherently aims at fostering positive social interactions. Moreover, our results show that attacking a specific component of the hate speech, and in particular its implicit negative stereotype and its hateful parts, leads to higher-quality generations.
2024: Helena Bonaldi, Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio, S. Villata, Marco Guerini
https://arxiv.org/pdf/2410.03466v1

… continue reading

298 episodes

Artwork

Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering

Papers Read on AI

33 subscribers

published 2M ago

Partager

MP3•Maison d'episode

Contenu fourni par Rob. Tout le contenu du podcast, y compris les épisodes, les graphiques et les descriptions de podcast, est téléchargé et fourni directement par Rob ou son partenaire de plateforme de podcast. Si vous pensez que quelqu'un utilise votre œuvre protégée sans votre autorisation, vous pouvez suivre le processus décrit ici https://fr.player.fm/legal.

The potential effectiveness of counterspeech as a hate speech mitigation strategy is attracting increasing interest in the NLG research community, particularly towards the task of automatically producing it. However, automatically generated responses often lack the argumentative richness which characterises expert-produced counterspeech. In this work, we focus on two aspects of counterspeech generation to produce more cogent responses. First, by investigating the tension between helpfulness and harmlessness of LLMs, we test whether the presence of safety guardrails hinders the quality of the generations. Secondly, we assess whether attacking a specific component of the hate speech results in a more effective argumentative strategy to fight online hate. By conducting an extensive human and automatic evaluation, we show how the presence of safety guardrails can be detrimental also to a task that inherently aims at fostering positive social interactions. Moreover, our results show that attacking a specific component of the hate speech, and in particular its implicit negative stereotype and its hateful parts, leads to higher-quality generations.
2024: Helena Bonaldi, Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio, S. Villata, Marco Guerini
https://arxiv.org/pdf/2410.03466v1

… continue reading

298 episodes

Tous les épisodes

×

Bienvenue sur Lecteur FM!

Lecteur FM recherche sur Internet des podcasts de haute qualité que vous pourrez apprécier dès maintenant. C'est la meilleure application de podcast et fonctionne sur Android, iPhone et le Web. Inscrivez-vous pour synchroniser les abonnements sur tous les appareils.

Écoutez plus de 500 sujets

Guide de référence rapide

Meilleurs podcasts

Les Grandes Gueules du Sport

Super Moscato Show

Sans rendez-vous - Mélanie Gomez

Entendez-vous l'éco ?

Les Grosses Têtes

Le Billet de Sophia Aram

M2 - Mutation numérique

La session de rattrapage, Jean-Luc Lemoine s’amuse de la télé

La Terre au carré

Sur les épaules de Darwin

Affaires sensibles

Hondelatte Raconte - Christophe Hondelatte

L'Heure Du Crime

CERNO L'anti-enquête

Scènes de Crime