Skip to the interactive examples
map of ideas

For researchers

Organize your reading.
Stay up to date.
Understand research.

We combine natural language processing, machine learning and citation analysis to find research relevant to your ideas and existing knowledge. Our analytical and AI tools help you understand what each paper contributes and how it connects to your work.

Explore example paper maps
Example workspace
About these examples
Real papers and saved results from the app.

Explore recorded searches, model analyses and paper maps. Actions here do not run new analyses or change your account.

The prerequisite example shows selected connections from a larger analysis. Each connection retains its model explanation.

Paper links and cited passages let you inspect the source evidence.

45 papers · 7 research areas
From attention to efficient inferenceCloser papers generally have more similar content.
Similarity: Sequence Transduction with Recurrent Neural Networks and Generating Sequences With Recurrent Neural NetworksSimilarity: Sequence Transduction with Recurrent Neural Networks and Sequence to Sequence Learning with Neural NetworksSimilarity: Sequence Transduction with Recurrent Neural Networks and Neural Machine Translation in Linear TimeSimilarity: Advances in Optimizing Recurrent Networks and Generating Sequences With Recurrent Neural NetworksSimilarity: Advances in Optimizing Recurrent Networks and Layer NormalizationSimilarity: Advances in Optimizing Recurrent Networks and AdaSplash: Adaptive Sparse Flash AttentionSimilarity: Maxout Networks and Deep Residual Learning for Image RecognitionSimilarity: Maxout Networks and Root Mean Square Layer NormalizationSimilarity: Maxout Networks and Reformer: The Efficient TransformerSimilarity: Maxout Networks and AdaSplash: Adaptive Sparse Flash AttentionSimilarity: Generating Sequences With Recurrent Neural Networks and Sequence to Sequence Learning with Neural NetworksSimilarity: Generating Sequences With Recurrent Neural Networks and Neural Machine Translation in Linear TimeSimilarity: Distributed Representations of Words and Phrases and their Compositionality and Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine TranslationSimilarity: Distributed Representations of Words and Phrases and their Compositionality and Strategies for Training Large Vocabulary Neural Language ModelsSimilarity: Distributed Representations of Words and Phrases and their Compositionality and Using the Output Embedding to Improve Language ModelsSimilarity: Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation and Neural Machine Translation by Jointly Learning to Align and TranslateSimilarity: Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation and On the Properties of Neural Machine Translation: Encoder-Decoder ApproachesSimilarity: Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation and Grammar as a Foreign LanguageSimilarity: Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation and Using the Output Embedding to Improve Language ModelsSimilarity: Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation and Neural Machine Translation in Linear TimeSimilarity: Neural Machine Translation by Jointly Learning to Align and Translate and Overcoming the Curse of Sentence Length for Neural Machine Translation using Automatic SegmentationSimilarity: Neural Machine Translation by Jointly Learning to Align and Translate and On the Properties of Neural Machine Translation: Encoder-Decoder ApproachesSimilarity: Neural Machine Translation by Jointly Learning to Align and Translate and Sequence to Sequence Learning with Neural NetworksSimilarity: Neural Machine Translation by Jointly Learning to Align and Translate and Effective Approaches to Attention-based Neural Machine TranslationSimilarity: Neural Machine Translation by Jointly Learning to Align and Translate and Neural Machine Translation in Linear TimeSimilarity: Neural Machine Translation by Jointly Learning to Align and Translate and Parallel Attention Mechanisms in Neural Machine TranslationSimilarity: Overcoming the Curse of Sentence Length for Neural Machine Translation using Automatic Segmentation and On the Properties of Neural Machine Translation: Encoder-Decoder ApproachesSimilarity: Overcoming the Curse of Sentence Length for Neural Machine Translation using Automatic Segmentation and Deep Architectures for Neural Machine TranslationSimilarity: On the Properties of Neural Machine Translation: Encoder-Decoder Approaches and Sequence to Sequence Learning with Neural NetworksSimilarity: On the Properties of Neural Machine Translation: Encoder-Decoder Approaches and Grammar as a Foreign LanguageSimilarity: On the Properties of Neural Machine Translation: Encoder-Decoder Approaches and Using the Output Embedding to Improve Language ModelsSimilarity: On the Properties of Neural Machine Translation: Encoder-Decoder Approaches and Neural Machine Translation in Linear TimeSimilarity: On the Properties of Neural Machine Translation: Encoder-Decoder Approaches and Parallel Attention Mechanisms in Neural Machine TranslationSimilarity: Sequence to Sequence Learning with Neural Networks and Neural Machine Translation in Linear TimeSimilarity: Grammar as a Foreign Language and Attention Is All You NeedSimilarity: Effective Approaches to Attention-based Neural Machine Translation and Attention Is All You NeedSimilarity: Effective Approaches to Attention-based Neural Machine Translation and Parallel Attention Mechanisms in Neural Machine TranslationSimilarity: Deep Residual Learning for Image Recognition and Root Mean Square Layer NormalizationSimilarity: Deep Residual Learning for Image Recognition and Reformer: The Efficient TransformerSimilarity: Strategies for Training Large Vocabulary Neural Language Models and Using the Output Embedding to Improve Language ModelsSimilarity: Strategies for Training Large Vocabulary Neural Language Models and AdaSplash: Adaptive Sparse Flash AttentionSimilarity: Strategies for Training Large Vocabulary Neural Language Models and DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference AccelerationSimilarity: Training Deep Nets with Sublinear Memory Cost and Layer NormalizationSimilarity: Training Deep Nets with Sublinear Memory Cost and Learning Fast Algorithms for Linear Transforms Using Butterfly FactorizationsSimilarity: Training Deep Nets with Sublinear Memory Cost and Data Movement Is All You Need: A Case Study on Optimizing TransformersSimilarity: Training Deep Nets with Sublinear Memory Cost and FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax AttentionSimilarity: Layer Normalization and Online normalizer calculation for softmaxSimilarity: Layer Normalization and Root Mean Square Layer NormalizationSimilarity: Using the Output Embedding to Improve Language Models and Neural Machine Translation in Linear TimeSimilarity: Neural Machine Translation in Linear Time and Attention Is All You NeedSimilarity: Attention Is All You Need and Deep Architectures for Neural Machine TranslationSimilarity: Attention Is All You Need and Parallel Attention Mechanisms in Neural Machine TranslationSimilarity: Attention Is All You Need and Efficient Inference For Neural Machine TranslationSimilarity: Deep Architectures for Neural Machine Translation and Parallel Attention Mechanisms in Neural Machine TranslationSimilarity: Deep Architectures for Neural Machine Translation and Efficient Inference For Neural Machine TranslationSimilarity: Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking and Kernel Operations on the GPU, with Autodiff, without Memory OverflowsSimilarity: Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking and Data Movement Is All You Need: A Case Study on Optimizing TransformersSimilarity: Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking and FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionSimilarity: Online normalizer calculation for softmax and Root Mean Square Layer NormalizationSimilarity: Online normalizer calculation for softmax and FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax AttentionSimilarity: Parallel Attention Mechanisms in Neural Machine Translation and Efficient Inference For Neural Machine TranslationSimilarity: Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations and Reformer: The Efficient TransformerSimilarity: Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations and Kernel Operations on the GPU, with Autodiff, without Memory OverflowsSimilarity: Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations and Data Movement Is All You Need: A Case Study on Optimizing TransformersSimilarity: Fast Transformer Decoding: One Write-Head is All You Need and Reformer: The Efficient TransformerSimilarity: Fast Transformer Decoding: One Write-Head is All You Need and FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessSimilarity: Fast Transformer Decoding: One Write-Head is All You Need and GQA: Training Generalized Multi-Query Transformer Models from Multi-Head CheckpointsSimilarity: Reformer: The Efficient Transformer and FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessSimilarity: Reformer: The Efficient Transformer and GQA: Training Generalized Multi-Query Transformer Models from Multi-Head CheckpointsSimilarity: Reformer: The Efficient Transformer and FlashMask: Efficient and Rich Mask Extension of FlashAttentionSimilarity: Kernel Operations on the GPU, with Autodiff, without Memory Overflows and Data Movement Is All You Need: A Case Study on Optimizing TransformersSimilarity: Data Movement Is All You Need: A Case Study on Optimizing Transformers and End-to-End Transformer Acceleration Through Processing-in-Memory ArchitecturesSimilarity: Data Movement Is All You Need: A Case Study on Optimizing Transformers and FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax AttentionSimilarity: Self-attention Does Not Need $O(n^2)$ Memory and FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessSimilarity: Self-attention Does Not Need $O(n^2)$ Memory and FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningSimilarity: Self-attention Does Not Need $O(n^2)$ Memory and Power Law Guided Dynamic Sifting for Efficient AttentionSimilarity: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness and An Exploration of Hierarchical Attention Transformers for Efficient Long Document ClassificationSimilarity: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness and FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningSimilarity: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness and FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionSimilarity: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness and FlashMask: Efficient and Rich Mask Extension of FlashAttentionSimilarity: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness and AdaSplash: Adaptive Sparse Flash AttentionSimilarity: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness and Power Law Guided Dynamic Sifting for Efficient AttentionSimilarity: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness and FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax AttentionSimilarity: An Exploration of Hierarchical Attention Transformers for Efficient Long Document Classification and AdaSplash: Adaptive Sparse Flash AttentionSimilarity: An Exploration of Hierarchical Attention Transformers for Efficient Long Document Classification and DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference AccelerationSimilarity: GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints and DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference AccelerationSimilarity: FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning and FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionSimilarity: FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning and FlashMask: Efficient and Rich Mask Extension of FlashAttentionSimilarity: FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning and SageAttention2++: A More Efficient Implementation of SageAttention2Similarity: FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision and FlashMask: Efficient and Rich Mask Extension of FlashAttentionSimilarity: FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision and SageAttention2++: A More Efficient Implementation of SageAttention2Similarity: FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision and Power Law Guided Dynamic Sifting for Efficient AttentionSimilarity: FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision and End-to-End Transformer Acceleration Through Processing-in-Memory ArchitecturesSimilarity: FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision and FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax AttentionSimilarity: FlashMask: Efficient and Rich Mask Extension of FlashAttention and SageAttention2++: A More Efficient Implementation of SageAttention2Similarity: FlashMask: Efficient and Rich Mask Extension of FlashAttention and FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax AttentionSimilarity: AdaSplash: Adaptive Sparse Flash Attention and RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM InferenceSimilarity: AdaSplash: Adaptive Sparse Flash Attention and DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference AccelerationSimilarity: AdaSplash: Adaptive Sparse Flash Attention and FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax AttentionSimilarity: RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference and DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference AccelerationSimilarity: RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference and HGCA: Hybrid GPU-CPU Attention for Long Context LLM InferenceSimilarity: Power Law Guided Dynamic Sifting for Efficient Attention and End-to-End Transformer Acceleration Through Processing-in-Memory ArchitecturesSimilarity: DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration and HGCA: Hybrid GPU-CPU Attention for Long Context LLM InferenceSimilarity: HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference and End-to-End Transformer Acceleration Through Processing-in-Memory ArchitecturesSequence Transduction with Recurrent Neural Networks · PaperAdvances in Optimizing Recurrent Networks · PaperMaxout Networks · PaperGenerating Sequences With Recurrent Neural Networks · PaperDistributed Representations of Words and Phrases and their Compositionality · PaperLearning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation · PaperNeural Machine Translation by Jointly Learning to Align and Translate · PaperOvercoming the Curse of Sentence Length for Neural Machine Translation using Automatic Segmentation · PaperOn the Properties of Neural Machine Translation: Encoder-Decoder Approaches · PaperSequence to Sequence Learning with Neural Networks · PaperGrammar as a Foreign Language · PaperEffective Approaches to Attention-based Neural Machine Translation · PaperDeep Residual Learning for Image Recognition · PaperStrategies for Training Large Vocabulary Neural Language Models · PaperTraining Deep Nets with Sublinear Memory Cost · PaperLayer Normalization · PaperUsing the Output Embedding to Improve Language Models · PaperNeural Machine Translation in Linear Time · PaperAttention Is All You Need · PaperDeep Architectures for Neural Machine Translation · PaperDissecting the NVIDIA Volta GPU Architecture via Microbenchmarking · PaperOnline normalizer calculation for softmax · PaperParallel Attention Mechanisms in Neural Machine Translation · PaperLearning Fast Algorithms for Linear Transforms Using Butterfly Factorizations · PaperRoot Mean Square Layer Normalization · PaperFast Transformer Decoding: One Write-Head is All You Need · PaperReformer: The Efficient Transformer · PaperKernel Operations on the GPU, with Autodiff, without Memory Overflows · PaperData Movement Is All You Need: A Case Study on Optimizing Transformers · PaperEfficient Inference For Neural Machine Translation · PaperSelf-attention Does Not Need $O(n^2)$ Memory · PaperFlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness · PaperAn Exploration of Hierarchical Attention Transformers for Efficient Long Document Classification · PaperGQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints · PaperFlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning · PaperFlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision · PaperFlashMask: Efficient and Rich Mask Extension of FlashAttention · PaperAdaSplash: Adaptive Sparse Flash Attention · PaperRetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference · PaperSageAttention2++: A More Efficient Implementation of SageAttention2 · PaperPower Law Guided Dynamic Sifting for Efficient Attention · PaperDAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration · PaperHGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference · PaperEnd-to-End Transformer Acceleration Through Processing-in-Memory Architectures · PaperFlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention · Paper
100%
Saved examples. Changes here last until you reload.Use your own papers

Free account

A free account keeps your papers, ideas and reading progress together. Follow up to two interests: each can be an idea or a set of papers.

New AI analyses require Connect. Paid subscriptions are not open yet; the saved examples above remain available to explore.
Go to your workspace