Research And Thesis
Interactive split-screen publication vault. Click any research paper in the catalog to inspect its deep findings, methodology, authors, and BibTeX citations live on the reader screen.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer et al. • Google Brain / Google Research
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril et al. • Meta AI
We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, showing that it is possible to train state-of-the-art models using publicly available datasets exclusively, without relying on proprietary and inaccessible datasets. In particular, LLaMA-13B outperforms GPT-3 (175B) on most benchmarks despite being more than 10x smaller.
In Search of an Understandable Consensus Algorithm (Raft)
Diego Ongaro, John Ousterhout • Stanford University
Raft is a consensus algorithm for managing a replicated log. It produces a result equivalent to (multi-)Paxos, and it is as efficient as Paxos, but its structure is different from Paxos; this makes Raft more understandable than Paxos and also provides a better foundation for building practical systems. Raft decomposes consensus into relatively independent subproblems: leader election, log replication, and safety.
Spanner: Google’s Globally-Distributed Database
James C. Corbett, Jeffrey Dean et al. • Google
Spanner is Google’s scalable, multi-version, globally-distributed, and synchronously-replicated database. It is the first system to distribute data at global scale and support externally-consistent distributed transactions. Spanner’s core innovation is TrueTime: an API that exposes clock uncertainty directly using GPS receivers and atomic clocks to enable serializable transactions without global locks.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao, Daniel Y. Fu et al. • Stanford University / Hazy Research
Transformers are slow and memory-hungry on long sequences, as the time and memory complexity of self-attention are quadratic in sequence length. We argue that a missing principle is making attention algorithm IO-aware—accounting for reads and writes between levels of GPU memory (SRAM and HBM). We propose FlashAttention, an IO-aware exact attention algorithm that uses tiling to reduce memory bandwidth overhead between GPU high-bandwidth memory (HBM) and fast SRAM cache.
Efficient and Robust Approximate Nearest Neighbor Search Using HNSW Graphs
Yu A. Malkov, D. A. Yashunin • Russian Academy of Sciences
Approximate nearest neighbor search in high dimensional spaces is a fundamental problem in data mining and vector search engines. We present the Hierarchical Navigable Small World (HNSW) graph algorithm, which builds multi-layer proximity graphs allowing logarithmic search complexity with high recall rates, serving as the core indexing technique in vector databases.
BBR: Congestion-Based Congestion Control
Neal Cardwell, Yuchung Cheng et al. • Google
TCP congestion control has historically interpreted packet loss as congestion. This paper introduces BBR (Bottleneck Bandwidth and RTT), an algorithm based on an explicit model of the network path that continuously measures the maximum bandwidth and minimum round-trip time, achieving dramatically higher throughput and lower latency across lossy connections.
A Survey of Zero-Knowledge Proofs in Decentralized Systems
Eli Ben-Sasson, Alessandro Chiesa et al. • Technion / UC Berkeley / Tel Aviv University
Zero-Knowledge Succinct Non-Interactive Arguments of Knowledge (ZK-SNARKs) enable a prover to convince a verifier that a mathematical statement is true without disclosing any secret witness. This paper analyzes polynomial commitment schemes, pairing-friendly elliptic curves, arithmetic circuits, and high-performance Rust prover architectures.