For researchers
Organize your reading.
Stay up to date.
Understand research.
We combine natural language processing, machine learning and citation analysis to find research relevant to your ideas and existing knowledge. Our analytical and AI tools help you understand what each paper contributes and how it connects to your work.
Faster exact attention for long documents
I want a Transformer to process longer documents without storing the full attention matrix. The computation should preserve exact attention while reducing reads and writes between GPU memory and on-chip storage. I would organize the work into blocks so that more GPU computation overlaps with data movement.
Has someone worked on a similar idea?
Use Find related papers to see research with related methods and goals.
Free account
A free account keeps your papers, ideas and reading progress together. Follow up to two interests: each can be an idea or a set of papers.
New AI analyses require Connect. Paid subscriptions are not open yet; the saved examples above remain available to explore.