Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention (AI Paper Explained)

Channel:

Yannic Kilcher

Subscribers:

300,000

Published on February 11, 2021 10:54:06 AM ● Video Link: https://www.youtube.com/watch?v=m-zrcmRd7E4

Duration: 48:12

15,953 views

528

#transformer #nystromer #nystromformer

The Nyströmformer (or Nystromformer, Nyströmer, Nystromer), is a new drop-in replacement for approximating the Self-Attention matrix in Transformers with linear memory and time requirements. Most importantly, it uses the Nystrom-Method to subselect (or segment mean) queries and keys as so-called landmarks and uses those to reconstruct the inherently low-rank attention matrix. This is relevant for many areas of Machine Learning, especially Natural Language processing, where it enables longer sequences of text to be processed at once.

OUTLINE:
0:00 - Intro & Overview
2:30 - The Quadratic Memory Bottleneck in Self-Attention
7:20 - The Softmax Operation in Attention
11:15 - Nyström-Approximation
14:00 - Getting Around the Softmax Problem
18:05 - Intuition for Landmark Method
28:05 - Full Algorithm
30:20 - Theoretical Guarantees
35:55 - Avoiding the Large Attention Matrix
36:55 - Subsampling Keys vs Negative Sampling
43:15 - Experimental Results
47:00 - Conclusion & Comments

Paper: https://arxiv.org/abs/2102.03902
Code: https://github.com/mlpen/Nystromformer
Appendix: https://github.com/mlpen/Nystromformer/blob/main/doc/Nystromformer_Supplement.pdf
LRA Results: https://twitter.com/tanmingxing/status/1359301186734620675
Twitter lucidrains w/ author: https://twitter.com/lucidrains/status/1359597104075661312
Twitter lucidrains w/ _clashluke: https://twitter.com/_clashluke/status/1359483460851802115

Abstract:
Transformers have emerged as a powerful tool for a broad range of natural language processing tasks. A key component that drives the impressive performance of Transformers is the self-attention mechanism that encodes the influence or dependence of other tokens on each specific token. While beneficial, the quadratic complexity of self-attention on the input sequence length has limited its application to longer sequences -- a topic being actively studied in the community. To address this limitation, we propose Nyströmformer -- a model that exhibits favorable scalability as a function of sequence length. Our idea is based on adapting the Nyström method to approximate standard self-attention with O(n) complexity. The scalability of Nyströmformer enables application to longer sequences with thousands of tokens. We perform evaluations on multiple downstream tasks on the GLUE benchmark and IMDB reviews with standard sequence length, and find that our Nyströmformer performs comparably, or in a few cases, even slightly better, than standard Transformer. Our code is at this https URL.

Authors: Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, Vikas Singh

Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yannic-kilcher
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/
BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

Other Videos By Yannic Kilcher

2021-03-16	Pretrained Transformers as Universal Computation Engines (Machine Learning Research Paper Explained)
2021-03-11	Yann LeCun - Self-Supervised Learning: The Dark Matter of Intelligence (FAIR Blog Post Explained)
2021-03-06	Apple or iPod??? Easy Fix for Adversarial Textual Attacks on OpenAI's CLIP Model! #Shorts
2021-03-05	Multimodal Neurons in Artificial Neural Networks (w/ OpenAI Microscope, Research Paper Explained)
2021-02-27	GLOM: How to represent part-whole hierarchies in a neural network (Geoff Hinton's Paper Explained)
2021-02-26	Linear Transformers Are Secretly Fast Weight Memory Systems (Machine Learning Paper Explained)
2021-02-25	DeBERTa: Decoding-enhanced BERT with Disentangled Attention (Machine Learning Paper Explained)
2021-02-19	Dreamer v2: Mastering Atari with Discrete World Models (Machine Learning Research Paper Explained)
2021-02-17	TransGAN: Two Transformers Can Make One Strong GAN (Machine Learning Research Paper Explained)
2021-02-14	NFNets: High-Performance Large-Scale Image Recognition Without Normalization (ML Paper Explained)
2021-02-11	Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention (AI Paper Explained)
2021-02-04	Deep Networks Are Kernel Machines (Paper Explained)
2021-02-02	Feedback Transformers: Addressing Some Limitations of Transformers with Feedback Memory (Explained)
2021-01-29	SingularityNET - A Decentralized, Open Market and Network for AIs (Whitepaper Explained)
2021-01-22	Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
2021-01-17	STOCHASTIC MEME DESCENT - Deep Learning Meme Review - Episode 2 (Part 2 of 2)
2021-01-12	OpenAI CLIP: ConnectingText and Images (Paper Explained)
2021-01-06	OpenAI DALL·E: Creating Images from Text (Blog Post Explained)
2020-12-26	Extracting Training Data from Large Language Models (Paper Explained)
2020-12-24	MEMES IS ALL YOU NEED - Deep Learning Meme Review - Episode 2 (Part 1 of 2)
2020-12-16	ReBeL - Combining Deep Reinforcement Learning and Search for Imperfect-Information Games (Explained)

Tags:

deep learning

machine learning

arxiv

explained

neural networks

artificial intelligence

paper

machine learning explained

transformers explained

nystrom

nystromformer

nystromer

nystrom approximation

self attention

attention mechanism

attention is all you need

transformer

linear transformer

linformer

linear attention

machine learning tutorial

quadratic attention

matrix approximation

low rank

landmark points

landmarks

matrix reconstruction

fast attention

Channel	Latest
Ini Guru Budi	6 hours ago
ALI NAWAZ ONLINE	6 hours ago
Viral Fun Aryan	6 hours ago
Mr Teddyy	6 hours ago
室友Cyo	6 hours ago
Otaku Beats	6 hours ago
Last Error Fixer	6 hours ago
X3pos	6 hours ago
AUTO KRYTYK	6 hours ago
VitoSinagaPrank	6 hours ago
Mettrox Live	6 hours ago
LCK	7 hours ago
* . ଘ buttercupxo ଓ . *	7 hours ago
Lil_murky20	7 hours ago
Chester	7 hours ago
GamerBill	7 hours ago
Đức Mạnh Melody	7 hours ago
Reo Gaming	7 hours ago
John Gage	7 hours ago
Rada	7 hours ago
DolmaKalem	7 hours ago
Momo Candy Gaming	7 hours ago
Left for Мыша	7 hours ago
ゲーム実況あおいさんとこ。	7 hours ago
Dota Kimono	7 hours ago