Artwork

Contenu fourni par Yannic Kilcher. Tout le contenu du podcast, y compris les épisodes, les graphiques et les descriptions de podcast, est téléchargé et fourni directement par Yannic Kilcher ou son partenaire de plateforme de podcast. Si vous pensez que quelqu'un utilise votre œuvre protégée sans votre autorisation, vous pouvez suivre le processus décrit ici https://fr.player.fm/legal.
Player FM - Application Podcast
Mettez-vous hors ligne avec l'application Player FM !

LAION-5B: 5 billion image-text-pairs dataset (with the authors)

58:01
 
Partager
 

Manage episode 326599118 series 2974171
Contenu fourni par Yannic Kilcher. Tout le contenu du podcast, y compris les épisodes, les graphiques et les descriptions de podcast, est téléchargé et fourni directement par Yannic Kilcher ou son partenaire de plateforme de podcast. Si vous pensez que quelqu'un utilise votre œuvre protégée sans votre autorisation, vous pouvez suivre le processus décrit ici https://fr.player.fm/legal.

#laion #clip #dalle

LAION-5B is an open, free dataset consisting of over 5 billion image-text-pairs. Today's video is an interview with three of its creators. We dive into the mechanics and challenges of operating at such large scale, how to keep cost low, what new possibilities are enabled with open datasets like this, and how to best handle safety and legal concerns.

OUTLINE:

0:00 - Intro

1:30 - Start of Interview

2:30 - What is LAION?

11:10 - What are the effects of CLIP filtering?

16:40 - How big is this dataset?

19:05 - Does the text always come from the alt-property?

22:45 - What does it take to work at scale?

25:50 -When will we replicate DALL-E?

31:30 - The surprisingly efficient pipeline

35:20 - How do you cover the S3 costs?

40:30 - Addressing safety & legal concerns

55:15 - Where can people get started?

References:

LAION website: https://laion.ai/

LAION Discord: https://discord.com/invite/mVcgxMPD7e

LAION-5B: https://laion.ai/laion-5b-a-new-era-o...

img2dataset tool: https://github.com/rom1504/img2dataset

LAION-400M: https://paperswithcode.com/dataset/la...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

  continue reading

177 episodes

Artwork
iconPartager
 
Manage episode 326599118 series 2974171
Contenu fourni par Yannic Kilcher. Tout le contenu du podcast, y compris les épisodes, les graphiques et les descriptions de podcast, est téléchargé et fourni directement par Yannic Kilcher ou son partenaire de plateforme de podcast. Si vous pensez que quelqu'un utilise votre œuvre protégée sans votre autorisation, vous pouvez suivre le processus décrit ici https://fr.player.fm/legal.

#laion #clip #dalle

LAION-5B is an open, free dataset consisting of over 5 billion image-text-pairs. Today's video is an interview with three of its creators. We dive into the mechanics and challenges of operating at such large scale, how to keep cost low, what new possibilities are enabled with open datasets like this, and how to best handle safety and legal concerns.

OUTLINE:

0:00 - Intro

1:30 - Start of Interview

2:30 - What is LAION?

11:10 - What are the effects of CLIP filtering?

16:40 - How big is this dataset?

19:05 - Does the text always come from the alt-property?

22:45 - What does it take to work at scale?

25:50 -When will we replicate DALL-E?

31:30 - The surprisingly efficient pipeline

35:20 - How do you cover the S3 costs?

40:30 - Addressing safety & legal concerns

55:15 - Where can people get started?

References:

LAION website: https://laion.ai/

LAION Discord: https://discord.com/invite/mVcgxMPD7e

LAION-5B: https://laion.ai/laion-5b-a-new-era-o...

img2dataset tool: https://github.com/rom1504/img2dataset

LAION-400M: https://paperswithcode.com/dataset/la...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

  continue reading

177 episodes

所有剧集

×
 
Loading …

Bienvenue sur Lecteur FM!

Lecteur FM recherche sur Internet des podcasts de haute qualité que vous pourrez apprécier dès maintenant. C'est la meilleure application de podcast et fonctionne sur Android, iPhone et le Web. Inscrivez-vous pour synchroniser les abonnements sur tous les appareils.

 

Guide de référence rapide