Home/Research

Research And Thesis

Interactive split-screen publication vault. Click any research paper in the catalog to inspect its deep findings, methodology, authors, and BibTeX citations live on the reader screen.

Active Reader ViewAI/MLPublished 2017
124,500 Citations

Attention Is All You Need

Institution: Google Brain / Google Research

Authors: Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin

Full Thesis Abstract & Findings:The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train.
Target Technologies & Architecture
PyTorchCUDATransformersPython
Research Catalog (8)Click any paper to inspect left
#1 Trending Thesis2017
Reading

Attention Is All You Need

Ashish Vaswani, Noam Shazeer et al. • Google Brain / Google Research

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train.

124,500 CitationsPyTorch
Up-to-date 2023–20242023
Inspect

LLaMA: Open and Efficient Foundation Language Models

Hugo Touvron, Thibaut Lavril et al. • Meta AI

We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, showing that it is possible to train state-of-the-art models using publicly available datasets exclusively, without relying on proprietary and inaccessible datasets. In particular, LLaMA-13B outperforms GPT-3 (175B) on most benchmarks despite being more than 10x smaller.

8,900 CitationsPyTorch
High-Impact Citation2014
Inspect

In Search of an Understandable Consensus Algorithm (Raft)

Diego Ongaro, John Ousterhout • Stanford University

Raft is a consensus algorithm for managing a replicated log. It produces a result equivalent to (multi-)Paxos, and it is as efficient as Paxos, but its structure is different from Paxos; this makes Raft more understandable than Paxos and also provides a better foundation for building practical systems. Raft decomposes consensus into relatively independent subproblems: leader election, log replication, and safety.

5,410 CitationsGo
High-Impact Citation2012
Inspect

Spanner: Google’s Globally-Distributed Database

James C. Corbett, Jeffrey Dean et al. • Google

Spanner is Google’s scalable, multi-version, globally-distributed, and synchronously-replicated database. It is the first system to distribute data at global scale and support externally-consistent distributed transactions. Spanner’s core innovation is TrueTime: an API that exposes clock uncertainty directly using GPS receivers and atomic clocks to enable serializable transactions without global locks.

4,120 CitationsC++
Peer-Reviewed2022
Inspect

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Tri Dao, Daniel Y. Fu et al. • Stanford University / Hazy Research

Transformers are slow and memory-hungry on long sequences, as the time and memory complexity of self-attention are quadratic in sequence length. We argue that a missing principle is making attention algorithm IO-aware—accounting for reads and writes between levels of GPU memory (SRAM and HBM). We propose FlashAttention, an IO-aware exact attention algorithm that uses tiling to reduce memory bandwidth overhead between GPU high-bandwidth memory (HBM) and fast SRAM cache.

3,820 CitationsCUDA
Peer-Reviewed2020
Inspect

Efficient and Robust Approximate Nearest Neighbor Search Using HNSW Graphs

Yu A. Malkov, D. A. Yashunin • Russian Academy of Sciences

Approximate nearest neighbor search in high dimensional spaces is a fundamental problem in data mining and vector search engines. We present the Hierarchical Navigable Small World (HNSW) graph algorithm, which builds multi-layer proximity graphs allowing logarithmic search complexity with high recall rates, serving as the core indexing technique in vector databases.

3,100 CitationsC++
Peer-Reviewed2016
Inspect

BBR: Congestion-Based Congestion Control

Neal Cardwell, Yuchung Cheng et al. • Google

TCP congestion control has historically interpreted packet loss as congestion. This paper introduces BBR (Bottleneck Bandwidth and RTT), an algorithm based on an explicit model of the network path that continuously measures the maximum bandwidth and minimum round-trip time, achieving dramatically higher throughput and lower latency across lossy connections.

2,200 CitationsLinux Kernel
Up-to-date 2023–20242023
Inspect

A Survey of Zero-Knowledge Proofs in Decentralized Systems

Eli Ben-Sasson, Alessandro Chiesa et al. • Technion / UC Berkeley / Tel Aviv University

Zero-Knowledge Succinct Non-Interactive Arguments of Knowledge (ZK-SNARKs) enable a prover to convince a verifier that a mathematical statement is true without disclosing any secret witness. This paper analyzes polynomial commitment schemes, pairing-friendly elliptic curves, arithmetic circuits, and high-performance Rust prover architectures.

1,100 CitationsRust
Sponsored Reference