0xDEADBEEF odkazy RSS | ⇉deadbeef blog⇉
Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors
A Brief Retrospective on SPARC Register Windows
Fast, small, robust: pick three. Introducing a novel branchless partition implementation.
Pattern-defeating Quicksort
Beautiful Branchless Binary Search
The quick and practical "MSI" hash table
Energy-efficient coarse-grain out-of-order execution
täko: A Polymorphic Cache Hierarchy for General-Purpose Optimization of Data Movement
Stuff your logs!
Closing the B-tree vs. LSM-tree Write Amplification Gap on Modern Storage Hardware with Built-in Transparent Compression
Applications on In-Order Multicore Architectures
Informing Memory Operations: Memory Performance Feedback Mechanisms and Their Applications
Discerning the Dominant Out-of-Order Performance Advantage: Is it Speculation or Dynamism?
iCFP: Tolerating All-Level Cache Misses in In-Order Processors
Dash: Scalable Hashing on Persistent Memory
Dashtable in Dragonfly
How fast are Linux pipes anyway?
Branch/cmove and compiler optimizations
Low-Latency, High-Throughput Garbage Collection
Performance Speed Limits
A whirlwind introduction to dataflow graphs
Recursive Linear Hashing
Dynamic Hashing Schemes
Changing std::sort at Google’s Scale and Beyond
Linear hashing: a newtool for file and table addressing
Entropy decoding in Oodle Data: Huffman decoding on the Jaguar
A linear streaming algorithm for community detection in very large networks
Towards Real-Time Community Detection in Large Networks
Metastable Failures in Distributed Systems
Memory Bandwidth Napkin Math
Static B-Trees
Useful old technologies: ASN.1
An Empirical Lower Bound on the Overheads of Production Garbage Collectors
Persistence for the Masses: RRB-Vectors in a Systems Language
Streaming Graph Partitioning for Large Distributed Graphs
Near linear time algorithm to detect community structures in large-scale networks
A Simple and Efficient Implementation for Small Databases
Popping the Hood on Golden Cove
AMX coprocessor
Evaluating the Cost of Atomic Operations onModern Architectures
Indirection Stream Semantic Register Architecture for Efficient Sparse-Dense Linear Algebra
Snitch: A tiny Pseudo Dual-Issue Processor for Area and Energy Efficient Execution of Floating-Point Intensive Workloads
Fibonacci Hashing: The Optimization that the World Forgot (or: a Better Alternative to Integer Modulo)
How Zen 2’s Op Cache Affects Performance
Better Bit Mixing - Improving on MurmurHash3's 64-bit Finalizer
Implementing Hash Tables in C
Auto-Predication of Critical Branches*
Evolution of the Samsung Exynos CPU Microarchitecture
Data Compression Accelerator on IBM POWER9and z15 Processors
Xuantie-910: A Commercial Multi-Core 12-Stage Pipeline Out-of-Order 64-bit High Performance RISC-V Processor with Vector Extension
What Is Macroscalar?
The Weird and Wacky World of VIA, the 3rd player in the “Modern” x86 market
Accelerating ML Recommendation with over a ThousandRISC-V/Tensor Processors on Esperanto’s ET-SoC-1 Chip
ARM or x86? ISA Doesn’t Matter
What scientists must know about hardware to write fast code
B-Trees: More Than I Thought I'd Want to Know
Don't Throw Out Your Algorithms Book Just Yet: Classical Data Structures That Can Outperform Learned Indexes
A fast alternative to the modulo reduction
Simple and Fast BlockQuicksort using Lomuto’s Partitioning Scheme
Accurate Throughput Prediction of Basic Blocks on Recent Intel Microarchitectures
Do Low-level Optimizations Matter?
shift_dfa.md
Beating the L1 cache with value speculation
Reverse-engineering the Mali G78
BlockQuicksort: How Branch Mispredictions don’t affect Quicksort
DeepWalk: Online Learning of Social Representations
Scaling Up All Pairs Similarity Search
LeapIO: Efficient and Portable Virtual NVMe Storageon ARM SoCs
ZoneFS - Zone filesystem for Zoned block devices
Don’t Be a Blockhead: Zoned Namespaces Make Workon Conventional SSDs Obsolete
Cores that don’t count
Computing the number of digits of an integer quickly
Hash, displace, and compress
External Memory Based Algorithm
I See Deadμops: Leaking Secrets via Intel/AMDMicro-Op Caches
Inheritance was invented as a performance hack
Apple M1: Load and Store Queue Measurements
FITing-Tree: A Data-aware Index Structure
An Efficient Algorithmfor Exploiting Multiple Arithmetic Units
IBM POWER9 processor core
Speculating the entire x86-64 Instruction Set In Seconds with This One Weird Trick
Apple M1 Microarchitecture Research (instruction latency and throughput)
WarpCore: A Library for fast Hash Tables on GPUs
A Variable Vector Length SIMD Architecture forHW/SW Co-designed Processors
KUTrace: Where have all the nanoseconds gone?
Benchmarking "Hello, World!"
Learned Garbage Collection
SLAP: A split latency adaptive VLIW pipeline architecture which enables on-the-fly variable SIMD vector-length
C-for-Metal: High Performance SIMD Programming on Intel GPUs
End-of-buffer checks in decompressors
NoFTL-KV: Tackling Write-Amplification on KV-Stores with Native Storage Management
Open-Channel SSD (What is it Good For)
Evolution of Development Priorities in Key-value Stores Serving Large-scale Applications:The RocksDB Experience
Zone Append: A New Way of Writing to Zoned Storage
The LibreSOC Project: Simple-V Vectorisation
Building Faster AMD64 Memset Routines
"RDNA 2" Instruction Set Architecture
But how, exactly, databases use mmap?
Inlining and Compiler Optimizations
That XOR Trick
Is this a branch?
Inline caching
ARM Cortex-A72 execution and load/store
On GPUs, ranges, latency, and superoptimisers
Parsing: a timeline
Generational References
A Comprehensive (and Animated) Guide to InnoDB Locking
RecSSD: Near Data Processing for Solid State Drive Based Recommendation Inference Extended Abstract
NOREBA: A Compiler-Informed Non-Speculative Out-of-Order Commit Processor Extended Abstract
SIMDRAM: A Framework for Bit-Serial SIMD Processing Using DRAM Extended Abstract
VEGEN: A Vectorizer Generator for SIMD and Beyond
Fast Local Page-Tables for Virtualized NUMA Servers with vMitosis Extended Abstract
AIR-FI:Generating Covert Wi-Fi Signals fromAir-Gapped Computers
Converting floating-point numbers to integers while preserving order
Regex literals optimization
D's Auto Decoding and You
A Rule-Based Style and Grammar Checker
How fast does interpolation search converge?
Transport triggered architecture
HiPEAC 2020 keynote 1: James Mickens on software-defined microarchitecture
Achieving 100Gbps intrusion prevention on a single server
PopCount on ARM64 in Go Assembler
Engineering In-place (Shared-memory) Sorting Algorithms
JIT Compiler of PCRE2
Producing Wrong Data Without Doing Anything Obviously Wrong!
An Empirical Evaluation of Set Similarity Join Techniques
An empirical evaluation of exact set similarity join techniques using GPUs
PM-LSH: A Fast and Accurate LSH Framework for High-Dimensional Approximate NN Search
Scalable Blocking for Very Large Databases
Finding Bytes in Arrays
A fastk-means implementation using coresets
Ridiculously fast unicode (UTF-8) validation
The Arm64 memory tagging extension in Linux
k-means++: The Advantages of Careful Seeding
A Fast Exact k-Nearest Neighbors Algorithm for High Dimensional Search Using k-Means Clustering and Triangle Inequality
Storage strategies for collections in dynamically typed languages
Why Aren’t More Users More Happy With Our VMs? Part 2
Why Aren’t More Users More Happy With Our VMs? Part 1
Custom Allocators Demystified
Loading CSV File at the Speed Limit of the NVMe Storage
When Network is Faster than Cache
Understand std::atomic::compare_exchange_weak() in C++11
SIMD transposes 1
D Slices
Static Analysis of Java Enterprise Applications: Frameworks and Caches, the Elephants in the Room
Hoare’s Rebuttal and Bubble Sort’s Comeback
HW and SW rules of thumb.
Faster intersections between sorted arrays with shotgun
Virtual Memory Tricks
CARAT: A Case for Virtual Memory through Compiler- and Runtime-Based Address Translation
The Cost of Software-Based Memory Management Without Virtual Memory
Go Your Own Way (Part Two: The Heap)
Go Your Own Way (Part One: The Stack)
Life in the Fast Lane
Don’t Fear the Reaper
Fast random pair divisive construction of kNN graph using generic distance measures
NN-Descent on High-Dimensional Data
Efficient K-Nearest Neighbor Graph Construction for Generic Similarity Measures
Improving Locality Sensitive Hashing by Efficiently Finding Projected Nearest Neighbors
On Modern Hardware the Min-Max Heap beats a Binary Heap
Sentinels can be faster
Performance Impact of Parallel Disk Access
split-forwarding.md
Zen 2 - Microarchitectures - AMD
Intelligent Probing for Locality Sensitive Hashing: Multi-Probe LSH and Beyond
Query-Aware Locality-Sensitive Hashing for Approximate Nearest Neighbor Search
Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting
Fast Search of Binary Codes with Distinctive Bits
Ultra Fast Medoid Identification via Correlated Sequential Halving
Medoids in almost linear time via multi-armed bandits
Fast Approximation of Centrality
fuzzy jaccard
The Power of Comparative Reasoning (WTA hash, winner take all)
What Is The Minimal Set Of Optimizations Needed For Zero-Cost Abstraction?
4K Aliasing
Garbage Collector Code Artifacts: Card Marking
Hardware Store Elimination
When Escape Analysis fails you?
Speculation in JavaScriptCore
The ABC’s of Templates in D
Coding for Random Projections
Min-Max Hash for Jaccard Similarity
Efficient nearest neighbors inspired by the fruit fly brain
Programmers Need To Learn Statistics Or I Will Kill Them All
SonicBOOM: The 3rd Generation Berkeley Out-of-Order Machine
LSH Forest: Self-Tuning Indexes for Similarity Search
Fast Intersection of Sorted Lists Using SSE Instructions
Latency implications of virtual memory
Why Java's TLABs are so important and why write contention is a performance killer in multicore environments
A Concurrency Cost Hierarchy
How JIT Compilers are Implemented and Fast: Julia, Pypy, LuaJIT, Graal and More
A Deep Introduction to JIT Compilers: JITs are not very Just-in-time
Improved Densification of One Permutation Hashing
Densifying One Permutation Hashing via Rotation for Fast Near Neighbor Search
Rapid Similarity Search with Weighted Min-Hash
'Fastware' - Andrei Alexandrescu
Moving Garbage Collection with Low-Variation Memory Overhead and Deterministic Concurrent Relocation
How do 'hot and cold' objects behave?
Advanced Matrix Extension (AMX) - x86
Asymmetric Minwise Hashing
Asymmetric LSH (ALSH) for Sublinear Time Maximum Inner Product Search (MIPS)
The rapid growth of io_uring
Radix sort: sorting integers (often) faster than std::sort.
Faster than radix sort: Kirkpatrick-Reisch sorting
The Cache Replacement Problem
The CHERI CPU Hardware software co design for security
Reflective Random Indexing and indirect inference: A scalable method for discovery of implicit connections
AVX-512 Mask Registers, Again
Benchmarking for Good with Aleksey Shipilev
Ice Lake Store Elimination
Lomuto’s Comeback
Unikernels: The Next Stage of Linux’s Dominance
How the Go runtime implements maps efficiently (without generics)
An history of NVidia Stream Multiprocessor
Graph-of-word and TW-IDF: New Approach to Ad Hoc IR
MIDAS: Microcluster-Based Detector of Anomalies in Edge Streams
Java Objects Inside Out
A Text Network Representation Model
Undocumented CPU Behavior: Analyzing Undocumented Opcodes on Intel x86-64
Understanding CPU Microarchitecture to Increase Performance
Measuring Time: From Java to Kernel and back
An empirical guide to the behavior and use of scalable persistent memory
Avoiding cache line overlap by replacing one 256-bit store with two 128-bit stores
Writing a full-text search engine using Bloom filters
Modern B-Tree Techniques
Everything I know about SSDs
A Look At Celerity’s Second-Gen 496-Core RISC-V Mesh NoC
Random Indexing Explained with High Probability
ComputeDRAM: In-Memory Compute Using Off-the-Shelf DRAMs
Performance variation in 2,386 ‘identical’ processors
A Position-Biased PageRank Algorithm for Keyphrase Extraction
TextRank: Bringing Order into Texts
Xor Filters: Faster and Smaller Than Bloom and CuckooFilters
Ambit: In-Memory Accelerator for Bulk Bitwise Operations Using Commodity DRAM Technology
Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives
Fast Bulk Bitwise AND and OR in DRAM
Computer Architecture - Lecture 6b: Computation in Memory I - Onur Mutlu
Base64 encoding and decoding at almost the speed of a memory copy
Could a Neuroscientist Understand a Microprocessor?
Hierarchical PLABs, CLABs, TLABs in Hotspot
Faster threshold queries with cache-sensitive scancount
Revec: Program Rejuvenation through Revectorization
Hiding Data in Hard-Drive’s Service Areas
Hard Drive of Hearing: Disks that Eavesdrop with a Synthesized Microphone
Future Directions for Optimizing Compilers
Multilayer ROP Protection via Microarchitectural Units Available in Commodity Hardware
x86-64 Instruction Usage among C/C++ Applications
Basic Performance Measurements of the Intel Optane DC Persistent Memory Module
Cheap transistors, expensive wires
A Hardware Accelerator for Tracing Garbage Collection
Word Hy-phen-a-tion by Com-put-er
I/O Is Faster Than the CPU – Let’s Partition Resources and Eliminate (Most) OS Abstractions
Design of the RISC-V Instruction Set Architecture
DIMMer: A case of turning off DIMMs in clouds.
And Then There Were None: A Stall-Free Real-Time Garbage Collector for Reconfigurable Hardware
Lost in translation: Exposing hidden compiler optimization opportunities
Accelerators for Data Processing
Array Bounds Check Elimination for the Java HotSpot TM Client Compiler
Mesh: Compacting Memory Management for C/C++ Applications
Large-Scale Reconfigurable Computing in a Microsoft Datacenter
The Case for Network-Accelerated Query Processing
Faster intersections between sorted arrays with shotgun
MinHashing
Why Systolic Architectures
BOLT: A Practical Binary Optimizer for Data Centers and Beyond
WaveFunctionCollapse (Bitmap & tilemap generation from a single example with the help of ideas from quantum mechanics.)
Beating hash tables with trees? The ART-ful radix trie
ISA Semantics for ARMv8-A, RISC-V, and CHERI-MIPS
Measuring the memory-level parallelism of a system using a small C++ program?
How to implement strings
How to Architect a Query Compiler, Revisited
Shenandoah GC: The Garbage Collector That Could : Aleksey Shipilev
Getting 4 bytes or a full cache line: same speed or not?
Dynamic Vectorization in the E2 Dynamic Multicore Architecture
An Evaluation of the TRIPS Computer System
Exploiting Superword Level Parallelism with Multimedia Instruction Sets
ispc: A SPMD Compiler for High-Performance CPU Programming
Lecture 7: The Programmable GPU Core
Mison: A Fast JSON Parser for Data Analytics
Active Pages: A Computation Model for Intelligent Memory
Radix Sort for Vector Multiprocessors
Scalable Processors in the Billion-Transistor Era: IRAM
Iterating in batches over data structures can be much faster…
ROLP: Runtime Object Lifetime Profiling for Big Data Memory Management
The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities
Fast, Flexible, Polyglot Instrumentation Support for Debuggers and other Tools
The surprising creativity of digital evolution
Parallel generational-copying garbage collection with a block-structured heap
JOP: A Java Optimized Processor
Using a Java Optimized Processor in a Real World Application
High-performance throughput computing
Fantastic Timers and Where to Find Them: High-Resolution Microarchitectural Attacks in JavaScript
What a difference a JVM makes?
JVM Anatomy Park #18: Scalar Replacement
Aleksandar Prokopec - Making Collection Operations Optimal with Aggressive JIT Compilation
KV-Direct: High-performance in-memory key-value store with programmable NIC
A fast alternative to the modulo reduction
Fast exact integer divisions using floating-point operations
UnifiedMap: How it works?
A Branchless UTF-8 Decoder
What every systems programmer should know about lockless concurrency
Zebras All the Way Down - Bryan Cantrill, Uptime 2017
Virtual Machine Warmup Blows Hot and Cold
The Renewed Case for the Reduced Instruction Set Computer: Avoiding ISA Bloat with Macro-Op Fusion for RISC-V
STREAM VBYTE: Faster Byte-Oriented Integer Compression
SFS: Random Write Considered Harmful in Solid State Drives
Energy Efficiency across Programming Languages
B-trees, Shadowing, and Clones
Strategies for Branch Target Buffers
The YAGS Branch Prediction Scheme
Dynamic Branch Prediction with Perceptrons
JavaScript for extending low-latency in-memory key-value stores
Efficient Immutable Collections
One-pass Code Generation in V8
Nom, a byte oriented, streaming, zero copy, parser combinators library in Rust
NOVA: A Log-structured File System for Hybrid Volatile/Non-volatile Main Memories
Breaking the x86 ISA
Pruning spaces from strings quickly on ARM processors
Top speed for top-k queries
Practical Partial Evaluation for High-Performance Dynamic Language Runtimes
NG2C: Pretenuring N-Generational GC for HotSpot Big Data Applications
QuickSelect versus binary heap for top-k queries
Quickly returning the top-k elements: computer science vs. the real world
Counting exactly the number of distinct elements: sorted arrays vs. hash sets?
Typed Architectures: architectural support for lightweight scripting
Adaptive Cuckoo Filters
Improving user perceived page load time using gaze
Glob Matching Can Be Simple And Fast Too
JIT and instanceof
Towards Efficient Dynamic Integer Overflow Detection on ARM Processors
Java Microbenchmark Harness: The Lesser of Two Evils
Beauty and the Burst: Remote Identification of Encrypted Video Streams
Vectorization in HotSpot JVM
Dynamic Coarse Grained Reconfigurable Architectures
The Structure and Performance of Efficient Interpreters
Strided Sampling Hashed Perceptron Predictor
Beyond the words: predicting user personality from heterogeneous information
Programming Languages: History and Future
Devirtualization
The Pauseless GC Algorithm
Grail Quest: A New Proposal for Hardware-assisted Garbage Collection
Tightly Packed Tries: How to Fit Large Models into Memory, and Make them Load Fast, Too
Multiprocessors Should Support Simple Memory Consistency Models
The Silently Shifting Semicolon
WeeFence: Toward Making Fences Free in TSO
Dynamo: A Transparent Dynamic Optimization System
Fast Haskell: Competing with C at parsing XML
Runtime Pointer Disambiguation
Be nice to your cache
DawnCC: a Source-to-Source Automatic Parallelizer of C and C++ Programs
Sorting improves word-aligned bitmap indexes
Faster Population Counts using AVX2 Instructions
Myths and Realities: The Performance Impact of Garbage Collection
The LuaJIT Wiki Garbage Collector
Transforming static data structures to dynamic structures
The Bloomier Filter: An Efficient Data Structure for Static Support Lookup Tables
Garbage Collection Algorithms
Scapegoat tree
Optimizing Data Structures in High-Level Programs - New Directions for Extensible Compilers based on Staging
FUNCTIONAL PEARL The Zipper
Cache Conscious Indexing for Decision-Support in Main Memory
Permutation Search Methods are Efficient, Yet Faster Search is Possible
Engineering Efficient and Effective Non-Metric Space Library
Succinct Nearest Neighbor Search (NAPP, Neighborhood approximation inverted index)
Learning to Prune in Metric and Non-Metric Spaces
Hamming Compatible Quantization for Hashing
Off the Beaten Path: Let’s Replace Term-Based Retrieval with k-NN Search
Large-Scale Distributed Locality-Sensitive Hashing for General Metric Data (DFLSH, Voroni LSH)
Effective Proximity Retrieval by Ordering Permutations
Speeding Up Permutation Based Indexing with Indexing
Practical and Optimal LSH for Angular Distance (cross-polytope LSH)
Large-scale similarity data management with distributed Metric Index (data space mapping)
Metric Space Searching Based on Random Bisectors and Binary Fingerprints
A Brief Index for Proximity Searching (Brief Permutation Index)
On Locality Sensitive Hashing in metric spaces (Brief Permutation Index)
LSH forest: self-tuning indexes for similarity search.
Multi-Probe LSH: Efficient Indexing for High-Dimensional Similarity Search
K-medoids LSH: a new locality sensitive hashing in general metric space
Comparative Analysis of Data Structures for Approximate Nearest Neighbor Search (small worls graphs are the fastest)
Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs
Accelerating Search and Recognition Workloads with SSE 4.2 String and Text Processing Instructions
Fast Sorted-Set Intersection using SIMD Instructions
ELB trees - efficient lock-free B+trees
Optimal Incremental Sorting
A Fast Algorithm for Computing Longest Common Subsequences
Algorithm for Computing Maximal Common Subsequence (longest common subsequence)
Ordered hash table
A Fast Write Barrier for Generational Garbage Collectors
SparseDTW: A Novel Approach to Speed up Dynamic Time Warping
SparseDTW: A Novel Approach to Speed up Dynamic Time Warping
FastDTW: Toward Accurate Dynamic Time Warping in Linear Time and Space
Efficient and Thread-Safe Objects for Dynamically-Typed Languages
How the JVM compares your strings using the craziest x86 instruction you've never heard of
Improving the Energy Efficiency of Big Cores
Why Nothing Matters: The Impact of Zeroing
Fast Algorithms for Sorting and Searching Strings (radix quicksort, multi-key quicksort)
Efficient Trie-Based Sorting of Large Sets of Strings
Burst Tries: A Fast, Efficient Data Structure for String Keys
Cache-Conscious Sorting of Large Sets of Strings with Dynamic Tries
The Average Case Analysis of Partition Sorts
Min-Max Heaps and Generalized Priority Queues
Partial Escape Analysis and Scalar Replacement for Java
Generic Multiset Programming with Discrimination-based Joins and Symbolic Cartesian Products
Generic Top-down Discrimination for Sorting and Partitioning in Linear Time
Bloofi: Multidimensional Bloom Filters
Lock Holder Preemption Avoidance via Transactional Lock Elision
Optimal sorting algorithms for parallel computers
Concurrent Search Tree by Lazy Splaying
Implementing sets efficiently in a functional language (trees of bounded balance)
Efficient Set Intersection for Inverted Indexing
Better bitmap performance with Roaring bitmaps
The Technology Behind Crusoe™ Processors
Simple, proven approaches to text retrieval
Virtually free - JVM callsite optimization by example
Programming Interfaces to Non-Volatile Memory
FlexSC: Flexible System Call Scheduling with Exception-Less System Calls
Fun C Micro-optimizations - restrict
Interlude: Numerical experiments in hashing
More numerical experiments in hashing: a conclusion (Robin Hood hashing)
Robin Hood Hashing should be your default Hash Table implementation
Robin Hood hashing
R-trees Have Grown Everywhere
Elastic Binary Trees - ebtree
Code Specialization for Memory Efficient Hash Tries (Short Paper)
Optimizing Hash-Array Mapped Tries for Fast and Lean Immutable JVM Collections
The Cache Performance and Optimization of Blocked Algorithms
Nearest neighbors and vector models – epilogue – curse of dimensionality
How does Java Both Optimize Hot Loops and Allow Debugging
Mobile Processors for Energy-Efficient Web Search
Finger Trees Custom Persistent Collections - Chris Houser
Evaluating HTM for pauseless garbage collectors in Java
Array layouts for comparison-based searching
Conc-Trees for Functional and Parallel Programming
High Performance and Scalable Radix Sorting: A case study of implementing dynamic parallelism for GPU computing
Fast sort on CPUs and GPUs: a case for bandwidth oblivious SIMD sort
Generic Top-down Discrimination for Sorting and Partitioning in Linear Time
Adaptive Lock-Free Maps: Purely-Functional to Scalable
Making Data Structures Persistent
Cache-oblivious data structures
Memory Coherence in Shared Virtual Memory Systems
Unsupervised Feature Selection on Data Streams / Streaming Anomaly Detection Using Randomized Matrix Sketching
ZuriHac 2015 - Discrimination is Wrong: Improving Productivity (Kmett)
IFL 2012. Fritz Henglein: Generic sorting and partitioning in linear time and fully abstractly
How Branch Mispredictions Affect Quicksort
PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort
Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited
Functional Pearl: A SQL to C Compiler in 500 Lines of Code
Fenwick Trees (prefix sums)
Branch Prediction and the Performance of Interpreters - Don't Trust Folklore
A Primer on Memory Consistency and Cache Coherence
Jump Threading
Faster Cover Trees
Carnegie Mellon - Parallel Computer Architecture 2012-Onur Mutlu - Lec 22 - Dataflow I
Adaptive Just-in-time Value Class Optimization
Trace-based Just-in-time Compilation for Lazy Functional Programming Languages
Adaptive Range Filters for Cold Data: Avoiding Trips to Siberia
Diff-Index: Differentiated Index in Distributed Log-Structured Data Stores
Stratified B-trees and versioning dictionaries
Log Structured Merge Trees (LSMT)
A Relational Model of Data for Large Shared Data Banks
Monet - A Next-Generation DBMS Kernel For Query-Intensive Applications
Your computer is already a distributed system. Why isn't your OS?
MonetDB/X100: Hyper-Pipelining Query Execution
MonetDB: Two Decades of Research in Column-oriented Database Architectures
mov is Turing-complete
SnapQueue: Lock-Free Queue with Constant Time Snapshots
Scalable Bloom Filters
Zero-Overhead Metaprogramming Reflection and Metaobject Protocols Fast and without Compromises
Early Experience with a Commercial Hardware Transactional Memory Implementation
Accelerating Native Calls using Transactional Memory
Architecture of a Database System
Purely Functional Data Structures (Okasaki)
CS 61B Lecture 34: Splay Trees
CS 61B Lecture 31: Disjoint Sets
CS 61B Lecture 36: Randomized Analysis
CS 61B Lecture 35: Amortized Analysis
Scalability! But at what COST? (paper)
Scalability! But at what COST?
IA Memory Ordering (x86 memory model)
Module 2.4 - Cache Coherence - 740: Computer Architecture 2013 - Carnegie Mellon - Onur Mutlu
Stable Distributions, Pseudorandom Generators, Embeddings, and Data Stream Computation
Hashing, sketching, and other approximate algorithms for high-dimensional data
Locality-Sensitive Hashing Scheme Based on p-Stable Distributions
Beyond the PDP-11: Architectural support for a memory-safe C abstract machine
Pycket: A Tracing JIT For a Functional Language
Collaborative Filtering Recommender Systems
Scaling Concurrent Log-Structured Data Stores
Asynchronized Concurrency: The Secret to Scaling Concurrent Search Data Structures
A Study of CRDTs that do Computations
Advanced Data Structures: Session 10: Dictionaries
The Design and Implementation of Modern Column-Oriented Database Systems
x86 is a high-level language
Introduction to HAMT
Sirius: An Open End-to-End Voice and Vision Personal Assistant and Its Implications for Future Warehouse Scale Computers
Staring into the Abyss: An Evaluation of Concurrency Control with One Thousand Cores
Paper: Staring into the Abyss: An Evaluation of Concurrency Control with One Thousand Cores
Compressed Text Indexes: From Theory to Practice!
Command-line tools can be 235x faster than your Hadoop cluster
Predecessor search for Big Data: x-fast tries, locality of reference and all that
Understanding and Expressing scalable Concurrency
The Art of Approximating Distributions: Histograms and Quantiles at Scale
q-digest - Medians and Beyond: New Aggregation Techniques for Sensor Networks
q-digest : an algorithm for computing approximate quantiles on a collection of integers
References for Data Stream Algorithms
A closer Look at GPUs
Lecture 14 - Out-of-Order Execution - Carnegie Mellon - Computer Architecture 2013 - Onur Mutlu
Operation fusion and deforestation for Scala
Specialized Evolution of the General-Purpose CPU
Immutability Changes Everything
The Design of Approximation Algorithms
Scalability! But at what COST?
Programming on Parallel Machines - GPU, Multicore, Clusters and More
Mergeable persistent data structures
High Performance Hardware-Accelerated Flash Key-Value Store
SipHash: a fast short-input PRF
Analysis of Pivot Sampling in Dual-Pivot Quicksort, A Holistic Analysis of Yaroslavskiy’s Partitioning Scheme
Suffix Trees and their Applications in String Algorithms
Borislav Petkov: x86 instruction encoding and the nasty hacks we do in the kernel
Don’t Thrash: How to Cache Your Hash on Flash
Resizable, Scalable, Concurrent Hash Tables via Relativistic Programming
Cuckoo Filter: Practically Better than Bloom
Resizable, Scalable, Concurrent Hash Tables via Relativistic Programming
What's the deal with Hardware Transactional Memory!?! (linux.conf.au 2014)
SimHash or the way to compare quickly two datasets
An Overview of Kernel Lock Improvements (on huge NUMA systems)
Invertible Bloom Lookup Tables
Instruction Sets Should Be Free: The Case For RISC-V
Accurate Methods for the Statistics of Surprise and Coincidence
Epiphany Architecture Reference
Perceptual Hashing (searching similar images, reverse image search)
The Bw-Tree: A B-tree for New Hardware Platforms
Philip Wadler: Why no one uses functional languages
Rank and select for succinct data structures
Bridging Islands of Specialized Code using Macros and Reified Types
Data types ala carte
Throw away the keys: Easy, Minimal Perfect Hashing
MICA: A Holistic Approach to Fast In-Memory Key-Value Storage
SQL versus coSQL — a compendium to Erik Meijer’s paper
Mark Hill CPU, TLB
Similarity Measurement on Leaf-labelled Trees
Modern Microprocessors - A 90 Minute Guide!
Is Parallel Programming Hard, And, If So, What Can You Do About It?, Hardware and its Habits
Why Functional Programming Matters
Fast Computation of min-Hash Signatures for Image Collections
Scalable, Example-Based Refactorings with Refaster
Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture
AA-Sort: A New Parallel Sorting Algorithm for Multi-Core SIMD Processors
An Experimental Study of Sorting and Branch Prediction
Data Structures and Algorithms for Nearest Neighbor Search in General Metric Spaces (VP-tree)
Duplicate News Story Detection Revisited
Monoids: Theme and Variations (Functional Pearl)
Amazon.com Recommendations - Item-to-Item Collaborative Filtering
HyperLogLog: the analysis of a near-optimal cardinality estimation algorithm
Phil Bagwell, Ideal Hash Trees, 2001
Phil Bagwell, Fast And Space Efficient Trie Searches, 2000
HyperLogLog in Practice: Algorithmic Engineering of a State of The Art Cardinality Estimation Algorithm
Memory system compression and its benefits
The Night Watch
Miniboxing: Improving the Speed to Code Size Tradeoff in Parametric Polymorphism Translations
Iterative Ranking from Pair-wise Comparisons
Fast Mergable Integer Maps
Sketch of the Day: Frugal Streaming (median, rank)
TRASH A dynamic LC-trie and hash data structure
RAY: Integrating Rx and Async for Direct-Style Reactive Streams
Instruction tables Lists of instruction latencies, throughputs and micro-operation break-downs for Intel, AMD and VIA CPUs
Memory Barriers: a Hardware View for Software Hackers
Monads for functional programming
RRB-Trees: Efficient Immutable Vectors
Five Myths about Hash Tables
dablooms - an open source, scalable, counting bloom filter library