0xDEADBEEF odkazy RSS | ⇉
deadbeef blog⇉
Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors A Brief Retrospective on SPARC Register Windows Fast, small, robust: pick three. Introducing a novel branchless partition implementation. Pattern-defeating Quicksort Beautiful Branchless Binary Search The quick and practical "MSI" hash table Energy-efficient coarse-grain out-of-order execution täko: A Polymorphic Cache Hierarchy for General-Purpose Optimization of Data Movement Stuff your logs! Closing the B-tree vs. LSM-tree Write Amplification Gap on Modern Storage Hardware with Built-in Transparent Compression Applications on In-Order Multicore Architectures Informing Memory Operations: Memory Performance Feedback Mechanisms and Their Applications Discerning the Dominant Out-of-Order Performance Advantage: Is it Speculation or Dynamism? iCFP: Tolerating All-Level Cache Misses in In-Order Processors Dash: Scalable Hashing on Persistent Memory Dashtable in Dragonfly How fast are Linux pipes anyway? Branch/cmove and compiler optimizations Low-Latency, High-Throughput Garbage Collection Performance Speed Limits A whirlwind introduction to dataflow graphs Recursive Linear Hashing Dynamic Hashing Schemes Changing std::sort at Google’s Scale and Beyond Linear hashing: a newtool for file and table addressing Entropy decoding in Oodle Data: Huffman decoding on the Jaguar A linear streaming algorithm for community detection in very large networks Towards Real-Time Community Detection in Large Networks Metastable Failures in Distributed Systems Memory Bandwidth Napkin Math Static B-Trees Useful old technologies: ASN.1 An Empirical Lower Bound on the Overheads of Production Garbage Collectors Persistence for the Masses: RRB-Vectors in a Systems Language Streaming Graph Partitioning for Large Distributed Graphs Near linear time algorithm to detect community structures in large-scale networks A Simple and Efficient Implementation for Small Databases Popping the Hood on Golden Cove AMX coprocessor Evaluating the Cost of Atomic Operations onModern Architectures Indirection Stream Semantic Register Architecture for Efficient Sparse-Dense Linear Algebra Snitch: A tiny Pseudo Dual-Issue Processor for Area and Energy Efficient Execution of Floating-Point Intensive Workloads Fibonacci Hashing: The Optimization that the World Forgot (or: a Better Alternative to Integer Modulo) How Zen 2’s Op Cache Affects Performance Better Bit Mixing - Improving on MurmurHash3's 64-bit Finalizer Implementing Hash Tables in C Auto-Predication of Critical Branches* Evolution of the Samsung Exynos CPU Microarchitecture Data Compression Accelerator on IBM POWER9and z15 Processors Xuantie-910: A Commercial Multi-Core 12-Stage Pipeline Out-of-Order 64-bit High Performance RISC-V Processor with Vector Extension What Is Macroscalar? The Weird and Wacky World of VIA, the 3rd player in the “Modern” x86 market Accelerating ML Recommendation with over a ThousandRISC-V/Tensor Processors on Esperanto’s ET-SoC-1 Chip ARM or x86? ISA Doesn’t Matter What scientists must know about hardware to write fast code B-Trees: More Than I Thought I'd Want to Know Don't Throw Out Your Algorithms Book Just Yet: Classical Data Structures That Can Outperform Learned Indexes A fast alternative to the modulo reduction Simple and Fast BlockQuicksort using Lomuto’s Partitioning Scheme Accurate Throughput Prediction of Basic Blocks on Recent Intel Microarchitectures Do Low-level Optimizations Matter? shift_dfa.md Beating the L1 cache with value speculation Reverse-engineering the Mali G78 BlockQuicksort: How Branch Mispredictions don’t affect Quicksort DeepWalk: Online Learning of Social Representations Scaling Up All Pairs Similarity Search LeapIO: Efficient and Portable Virtual NVMe Storageon ARM SoCs ZoneFS - Zone filesystem for Zoned block devices Don’t Be a Blockhead: Zoned Namespaces Make Workon Conventional SSDs Obsolete Cores that don’t count Computing the number of digits of an integer quickly Hash, displace, and compress External Memory Based Algorithm I See Deadμops: Leaking Secrets via Intel/AMDMicro-Op Caches Inheritance was invented as a performance hack Apple M1: Load and Store Queue Measurements FITing-Tree: A Data-aware Index Structure An Efficient Algorithmfor Exploiting Multiple Arithmetic Units IBM POWER9 processor core Speculating the entire x86-64 Instruction Set In Seconds with This One Weird Trick Apple M1 Microarchitecture Research (instruction latency and throughput) WarpCore: A Library for fast Hash Tables on GPUs A Variable Vector Length SIMD Architecture forHW/SW Co-designed Processors KUTrace: Where have all the nanoseconds gone? Benchmarking "Hello, World!" Learned Garbage Collection SLAP: A split latency adaptive VLIW pipeline architecture which enables on-the-fly variable SIMD vector-length C-for-Metal: High Performance SIMD Programming on Intel GPUs End-of-buffer checks in decompressors NoFTL-KV: Tackling Write-Amplification on KV-Stores with Native Storage Management Open-Channel SSD (What is it Good For) Evolution of Development Priorities in Key-value Stores Serving Large-scale Applications:The RocksDB Experience Zone Append: A New Way of Writing to Zoned Storage The LibreSOC Project: Simple-V Vectorisation Building Faster AMD64 Memset Routines "RDNA 2" Instruction Set Architecture But how, exactly, databases use mmap? Inlining and Compiler Optimizations That XOR Trick Is this a branch? Inline caching ARM Cortex-A72 execution and load/store On GPUs, ranges, latency, and superoptimisers Parsing: a timeline Generational References A Comprehensive (and Animated) Guide to InnoDB Locking RecSSD: Near Data Processing for Solid State Drive Based Recommendation Inference Extended Abstract NOREBA: A Compiler-Informed Non-Speculative Out-of-Order Commit Processor Extended Abstract SIMDRAM: A Framework for Bit-Serial SIMD Processing Using DRAM Extended Abstract VEGEN: A Vectorizer Generator for SIMD and Beyond Fast Local Page-Tables for Virtualized NUMA Servers with vMitosis Extended Abstract AIR-FI:Generating Covert Wi-Fi Signals fromAir-Gapped Computers Converting floating-point numbers to integers while preserving order Regex literals optimization D's Auto Decoding and You A Rule-Based Style and Grammar Checker How fast does interpolation search converge? Transport triggered architecture HiPEAC 2020 keynote 1: James Mickens on software-defined microarchitecture Achieving 100Gbps intrusion prevention on a single server PopCount on ARM64 in Go Assembler Engineering In-place (Shared-memory) Sorting Algorithms JIT Compiler of PCRE2 Producing Wrong Data Without Doing Anything Obviously Wrong! An Empirical Evaluation of Set Similarity Join Techniques An empirical evaluation of exact set similarity join techniques using GPUs PM-LSH: A Fast and Accurate LSH Framework for High-Dimensional Approximate NN Search Scalable Blocking for Very Large Databases Finding Bytes in Arrays A fastk-means implementation using coresets Ridiculously fast unicode (UTF-8) validation The Arm64 memory tagging extension in Linux k-means++: The Advantages of Careful Seeding A Fast Exact k-Nearest Neighbors Algorithm for High Dimensional Search Using k-Means Clustering and Triangle Inequality Storage strategies for collections in dynamically typed languages Why Aren’t More Users More Happy With Our VMs? Part 2 Why Aren’t More Users More Happy With Our VMs? Part 1 Custom Allocators Demystified Loading CSV File at the Speed Limit of the NVMe Storage When Network is Faster than Cache Understand std::atomic::compare_exchange_weak() in C++11 SIMD transposes 1 D Slices Static Analysis of Java Enterprise Applications: Frameworks and Caches, the Elephants in the Room Hoare’s Rebuttal and Bubble Sort’s Comeback HW and SW rules of thumb. Faster intersections between sorted arrays with shotgun Virtual Memory Tricks CARAT: A Case for Virtual Memory through Compiler- and Runtime-Based Address Translation The Cost of Software-Based Memory Management Without Virtual Memory Go Your Own Way (Part Two: The Heap) Go Your Own Way (Part One: The Stack) Life in the Fast Lane Don’t Fear the Reaper Fast random pair divisive construction of kNN graph using generic distance measures NN-Descent on High-Dimensional Data Efficient K-Nearest Neighbor Graph Construction for Generic Similarity Measures Improving Locality Sensitive Hashing by Efficiently Finding Projected Nearest Neighbors On Modern Hardware the Min-Max Heap beats a Binary Heap Sentinels can be faster Performance Impact of Parallel Disk Access split-forwarding.md Zen 2 - Microarchitectures - AMD Intelligent Probing for Locality Sensitive Hashing: Multi-Probe LSH and Beyond Query-Aware Locality-Sensitive Hashing for Approximate Nearest Neighbor Search Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting Fast Search of Binary Codes with Distinctive Bits Ultra Fast Medoid Identification via Correlated Sequential Halving Medoids in almost linear time via multi-armed bandits Fast Approximation of Centrality fuzzy jaccard The Power of Comparative Reasoning (WTA hash, winner take all) What Is The Minimal Set Of Optimizations Needed For Zero-Cost Abstraction? 4K Aliasing Garbage Collector Code Artifacts: Card Marking Hardware Store Elimination When Escape Analysis fails you? Speculation in JavaScriptCore The ABC’s of Templates in D Coding for Random Projections Min-Max Hash for Jaccard Similarity Efficient nearest neighbors inspired by the fruit fly brain Programmers Need To Learn Statistics Or I Will Kill Them All SonicBOOM: The 3rd Generation Berkeley Out-of-Order Machine LSH Forest: Self-Tuning Indexes for Similarity Search Fast Intersection of Sorted Lists Using SSE Instructions Latency implications of virtual memory Why Java's TLABs are so important and why write contention is a performance killer in multicore environments A Concurrency Cost Hierarchy How JIT Compilers are Implemented and Fast: Julia, Pypy, LuaJIT, Graal and More A Deep Introduction to JIT Compilers: JITs are not very Just-in-time Improved Densification of One Permutation Hashing Densifying One Permutation Hashing via Rotation for Fast Near Neighbor Search Rapid Similarity Search with Weighted Min-Hash 'Fastware' - Andrei Alexandrescu Moving Garbage Collection with Low-Variation Memory Overhead and Deterministic Concurrent Relocation How do 'hot and cold' objects behave? Advanced Matrix Extension (AMX) - x86 Asymmetric Minwise Hashing Asymmetric LSH (ALSH) for Sublinear Time Maximum Inner Product Search (MIPS) The rapid growth of io_uring Radix sort: sorting integers (often) faster than std::sort. Faster than radix sort: Kirkpatrick-Reisch sorting The Cache Replacement Problem The CHERI CPU Hardware software co design for security Reflective Random Indexing and indirect inference: A scalable method for discovery of implicit connections AVX-512 Mask Registers, Again Benchmarking for Good with Aleksey Shipilev Ice Lake Store Elimination Lomuto’s Comeback Unikernels: The Next Stage of Linux’s Dominance How the Go runtime implements maps efficiently (without generics) An history of NVidia Stream Multiprocessor Graph-of-word and TW-IDF: New Approach to Ad Hoc IR MIDAS: Microcluster-Based Detector of Anomalies in Edge Streams Java Objects Inside Out A Text Network Representation Model Undocumented CPU Behavior: Analyzing Undocumented Opcodes on Intel x86-64 Understanding CPU Microarchitecture to Increase Performance Measuring Time: From Java to Kernel and back An empirical guide to the behavior and use of scalable persistent memory Avoiding cache line overlap by replacing one 256-bit store with two 128-bit stores Writing a full-text search engine using Bloom filters Modern B-Tree Techniques Everything I know about SSDs A Look At Celerity’s Second-Gen 496-Core RISC-V Mesh NoC Random Indexing Explained with High Probability ComputeDRAM: In-Memory Compute Using Off-the-Shelf DRAMs Performance variation in 2,386 ‘identical’ processors A Position-Biased PageRank Algorithm for Keyphrase Extraction TextRank: Bringing Order into Texts Xor Filters: Faster and Smaller Than Bloom and CuckooFilters Ambit: In-Memory Accelerator for Bulk Bitwise Operations Using Commodity DRAM Technology Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives Fast Bulk Bitwise AND and OR in DRAM Computer Architecture - Lecture 6b: Computation in Memory I - Onur Mutlu Base64 encoding and decoding at almost the speed of a memory copy Could a Neuroscientist Understand a Microprocessor? Hierarchical PLABs, CLABs, TLABs in Hotspot Faster threshold queries with cache-sensitive scancount Revec: Program Rejuvenation through Revectorization Hiding Data in Hard-Drive’s Service Areas Hard Drive of Hearing: Disks that Eavesdrop with a Synthesized Microphone Future Directions for Optimizing Compilers Multilayer ROP Protection via Microarchitectural Units Available in Commodity Hardware x86-64 Instruction Usage among C/C++ Applications Basic Performance Measurements of the Intel Optane DC Persistent Memory Module Cheap transistors, expensive wires A Hardware Accelerator for Tracing Garbage Collection Word Hy-phen-a-tion by Com-put-er I/O Is Faster Than the CPU – Let’s Partition Resources and Eliminate (Most) OS Abstractions Design of the RISC-V Instruction Set Architecture DIMMer: A case of turning off DIMMs in clouds. And Then There Were None: A Stall-Free Real-Time Garbage Collector for Reconfigurable Hardware Lost in translation: Exposing hidden compiler optimization opportunities Accelerators for Data Processing Array Bounds Check Elimination for the Java HotSpot TM Client Compiler Mesh: Compacting Memory Management for C/C++ Applications Large-Scale Reconfigurable Computing in a Microsoft Datacenter The Case for Network-Accelerated Query Processing Faster intersections between sorted arrays with shotgun MinHashing Why Systolic Architectures BOLT: A Practical Binary Optimizer for Data Centers and Beyond WaveFunctionCollapse (Bitmap & tilemap generation from a single example with the help of ideas from quantum mechanics.) Beating hash tables with trees? The ART-ful radix trie ISA Semantics for ARMv8-A, RISC-V, and CHERI-MIPS Measuring the memory-level parallelism of a system using a small C++ program? How to implement strings How to Architect a Query Compiler, Revisited Shenandoah GC: The Garbage Collector That Could : Aleksey Shipilev Getting 4 bytes or a full cache line: same speed or not? Dynamic Vectorization in the E2 Dynamic Multicore Architecture An Evaluation of the TRIPS Computer System Exploiting Superword Level Parallelism with Multimedia Instruction Sets ispc: A SPMD Compiler for High-Performance CPU Programming Lecture 7: The Programmable GPU Core Mison: A Fast JSON Parser for Data Analytics Active Pages: A Computation Model for Intelligent Memory Radix Sort for Vector Multiprocessors Scalable Processors in the Billion-Transistor Era: IRAM Iterating in batches over data structures can be much faster… ROLP: Runtime Object Lifetime Profiling for Big Data Memory Management The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities Fast, Flexible, Polyglot Instrumentation Support for Debuggers and other Tools The surprising creativity of digital evolution Parallel generational-copying garbage collection with a block-structured heap JOP: A Java Optimized Processor Using a Java Optimized Processor in a Real World Application High-performance throughput computing Fantastic Timers and Where to Find Them: High-Resolution Microarchitectural Attacks in JavaScript What a difference a JVM makes? JVM Anatomy Park #18: Scalar Replacement Aleksandar Prokopec - Making Collection Operations Optimal with Aggressive JIT Compilation KV-Direct: High-performance in-memory key-value store with programmable NIC A fast alternative to the modulo reduction Fast exact integer divisions using floating-point operations UnifiedMap: How it works? A Branchless UTF-8 Decoder What every systems programmer should know about lockless concurrency Zebras All the Way Down - Bryan Cantrill, Uptime 2017 Virtual Machine Warmup Blows Hot and Cold The Renewed Case for the Reduced Instruction Set Computer: Avoiding ISA Bloat with Macro-Op Fusion for RISC-V STREAM VBYTE: Faster Byte-Oriented Integer Compression SFS: Random Write Considered Harmful in Solid State Drives Energy Efficiency across Programming Languages B-trees, Shadowing, and Clones Strategies for Branch Target Buffers The YAGS Branch Prediction Scheme Dynamic Branch Prediction with Perceptrons JavaScript for extending low-latency in-memory key-value stores Efficient Immutable Collections One-pass Code Generation in V8 Nom, a byte oriented, streaming, zero copy, parser combinators library in Rust NOVA: A Log-structured File System for Hybrid Volatile/Non-volatile Main Memories Breaking the x86 ISA Pruning spaces from strings quickly on ARM processors Top speed for top-k queries Practical Partial Evaluation for High-Performance Dynamic Language Runtimes NG2C: Pretenuring N-Generational GC for HotSpot Big Data Applications QuickSelect versus binary heap for top-k queries Quickly returning the top-k elements: computer science vs. the real world Counting exactly the number of distinct elements: sorted arrays vs. hash sets? Typed Architectures: architectural support for lightweight scripting Adaptive Cuckoo Filters Improving user perceived page load time using gaze Glob Matching Can Be Simple And Fast Too JIT and instanceof Towards Efficient Dynamic Integer Overflow Detection on ARM Processors Java Microbenchmark Harness: The Lesser of Two Evils Beauty and the Burst: Remote Identification of Encrypted Video Streams Vectorization in HotSpot JVM Dynamic Coarse Grained Reconfigurable Architectures The Structure and Performance of Efficient Interpreters Strided Sampling Hashed Perceptron Predictor Beyond the words: predicting user personality from heterogeneous information Programming Languages: History and Future Devirtualization The Pauseless GC Algorithm Grail Quest: A New Proposal for Hardware-assisted Garbage Collection Tightly Packed Tries: How to Fit Large Models into Memory, and Make them Load Fast, Too Multiprocessors Should Support Simple Memory Consistency Models The Silently Shifting Semicolon WeeFence: Toward Making Fences Free in TSO Dynamo: A Transparent Dynamic Optimization System Fast Haskell: Competing with C at parsing XML Runtime Pointer Disambiguation Be nice to your cache DawnCC: a Source-to-Source Automatic Parallelizer of C and C++ Programs Sorting improves word-aligned bitmap indexes Faster Population Counts using AVX2 Instructions Myths and Realities: The Performance Impact of Garbage Collection The LuaJIT Wiki Garbage Collector Transforming static data structures to dynamic structures The Bloomier Filter: An Efficient Data Structure for Static Support Lookup Tables Garbage Collection Algorithms Scapegoat tree Optimizing Data Structures in High-Level Programs - New Directions for Extensible Compilers based on Staging FUNCTIONAL PEARL The Zipper Cache Conscious Indexing for Decision-Support in Main Memory Permutation Search Methods are Efficient, Yet Faster Search is Possible Engineering Efficient and Effective Non-Metric Space Library Succinct Nearest Neighbor Search (NAPP, Neighborhood approximation inverted index) Learning to Prune in Metric and Non-Metric Spaces Hamming Compatible Quantization for Hashing Off the Beaten Path: Let’s Replace Term-Based Retrieval with k-NN Search Large-Scale Distributed Locality-Sensitive Hashing for General Metric Data (DFLSH, Voroni LSH) Effective Proximity Retrieval by Ordering Permutations Speeding Up Permutation Based Indexing with Indexing Practical and Optimal LSH for Angular Distance (cross-polytope LSH) Large-scale similarity data management with distributed Metric Index (data space mapping) Metric Space Searching Based on Random Bisectors and Binary Fingerprints A Brief Index for Proximity Searching (Brief Permutation Index) On Locality Sensitive Hashing in metric spaces (Brief Permutation Index) LSH forest: self-tuning indexes for similarity search. Multi-Probe LSH: Efficient Indexing for High-Dimensional Similarity Search K-medoids LSH: a new locality sensitive hashing in general metric space Comparative Analysis of Data Structures for Approximate Nearest Neighbor Search (small worls graphs are the fastest) Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs Accelerating Search and Recognition Workloads with SSE 4.2 String and Text Processing Instructions Fast Sorted-Set Intersection using SIMD Instructions ELB trees - efficient lock-free B+trees Optimal Incremental Sorting A Fast Algorithm for Computing Longest Common Subsequences Algorithm for Computing Maximal Common Subsequence (longest common subsequence) Ordered hash table A Fast Write Barrier for Generational Garbage Collectors SparseDTW: A Novel Approach to Speed up Dynamic Time Warping SparseDTW: A Novel Approach to Speed up Dynamic Time Warping FastDTW: Toward Accurate Dynamic Time Warping in Linear Time and Space Efficient and Thread-Safe Objects for Dynamically-Typed Languages How the JVM compares your strings using the craziest x86 instruction you've never heard of Improving the Energy Efficiency of Big Cores Why Nothing Matters: The Impact of Zeroing Fast Algorithms for Sorting and Searching Strings (radix quicksort, multi-key quicksort) Efficient Trie-Based Sorting of Large Sets of Strings Burst Tries: A Fast, Efficient Data Structure for String Keys Cache-Conscious Sorting of Large Sets of Strings with Dynamic Tries The Average Case Analysis of Partition Sorts Min-Max Heaps and Generalized Priority Queues Partial Escape Analysis and Scalar Replacement for Java Generic Multiset Programming with Discrimination-based Joins and Symbolic Cartesian Products Generic Top-down Discrimination for Sorting and Partitioning in Linear Time Bloofi: Multidimensional Bloom Filters Lock Holder Preemption Avoidance via Transactional Lock Elision Optimal sorting algorithms for parallel computers Concurrent Search Tree by Lazy Splaying Implementing sets efficiently in a functional language (trees of bounded balance) Efficient Set Intersection for Inverted Indexing Better bitmap performance with Roaring bitmaps The Technology Behind Crusoe™ Processors Simple, proven approaches to text retrieval Virtually free - JVM callsite optimization by example Programming Interfaces to Non-Volatile Memory FlexSC: Flexible System Call Scheduling with Exception-Less System Calls Fun C Micro-optimizations - restrict Interlude: Numerical experiments in hashing More numerical experiments in hashing: a conclusion (Robin Hood hashing) Robin Hood Hashing should be your default Hash Table implementation Robin Hood hashing R-trees Have Grown Everywhere Elastic Binary Trees - ebtree Code Specialization for Memory Efficient Hash Tries (Short Paper) Optimizing Hash-Array Mapped Tries for Fast and Lean Immutable JVM Collections The Cache Performance and Optimization of Blocked Algorithms Nearest neighbors and vector models – epilogue – curse of dimensionality How does Java Both Optimize Hot Loops and Allow Debugging Mobile Processors for Energy-Efficient Web Search Finger Trees Custom Persistent Collections - Chris Houser Evaluating HTM for pauseless garbage collectors in Java Array layouts for comparison-based searching Conc-Trees for Functional and Parallel Programming High Performance and Scalable Radix Sorting: A case study of implementing dynamic parallelism for GPU computing Fast sort on CPUs and GPUs: a case for bandwidth oblivious SIMD sort Generic Top-down Discrimination for Sorting and Partitioning in Linear Time Adaptive Lock-Free Maps: Purely-Functional to Scalable Making Data Structures Persistent Cache-oblivious data structures Memory Coherence in Shared Virtual Memory Systems Unsupervised Feature Selection on Data Streams / Streaming Anomaly Detection Using Randomized Matrix Sketching ZuriHac 2015 - Discrimination is Wrong: Improving Productivity (Kmett) IFL 2012. Fritz Henglein: Generic sorting and partitioning in linear time and fully abstractly How Branch Mispredictions Affect Quicksort PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited Functional Pearl: A SQL to C Compiler in 500 Lines of Code Fenwick Trees (prefix sums) Branch Prediction and the Performance of Interpreters - Don't Trust Folklore A Primer on Memory Consistency and Cache Coherence Jump Threading Faster Cover Trees Carnegie Mellon - Parallel Computer Architecture 2012-Onur Mutlu - Lec 22 - Dataflow I Adaptive Just-in-time Value Class Optimization Trace-based Just-in-time Compilation for Lazy Functional Programming Languages Adaptive Range Filters for Cold Data: Avoiding Trips to Siberia Diff-Index: Differentiated Index in Distributed Log-Structured Data Stores Stratified B-trees and versioning dictionaries Log Structured Merge Trees (LSMT) A Relational Model of Data for Large Shared Data Banks Monet - A Next-Generation DBMS Kernel For Query-Intensive Applications Your computer is already a distributed system. Why isn't your OS? MonetDB/X100: Hyper-Pipelining Query Execution MonetDB: Two Decades of Research in Column-oriented Database Architectures mov is Turing-complete SnapQueue: Lock-Free Queue with Constant Time Snapshots Scalable Bloom Filters Zero-Overhead Metaprogramming Reflection and Metaobject Protocols Fast and without Compromises Early Experience with a Commercial Hardware Transactional Memory Implementation Accelerating Native Calls using Transactional Memory Architecture of a Database System Purely Functional Data Structures (Okasaki) CS 61B Lecture 34: Splay Trees CS 61B Lecture 31: Disjoint Sets CS 61B Lecture 36: Randomized Analysis CS 61B Lecture 35: Amortized Analysis Scalability! But at what COST? (paper) Scalability! But at what COST? IA Memory Ordering (x86 memory model) Module 2.4 - Cache Coherence - 740: Computer Architecture 2013 - Carnegie Mellon - Onur Mutlu Stable Distributions, Pseudorandom Generators, Embeddings, and Data Stream Computation Hashing, sketching, and other approximate algorithms for high-dimensional data Locality-Sensitive Hashing Scheme Based on p-Stable Distributions Beyond the PDP-11: Architectural support for a memory-safe C abstract machine Pycket: A Tracing JIT For a Functional Language Collaborative Filtering Recommender Systems Scaling Concurrent Log-Structured Data Stores Asynchronized Concurrency: The Secret to Scaling Concurrent Search Data Structures A Study of CRDTs that do Computations Advanced Data Structures: Session 10: Dictionaries The Design and Implementation of Modern Column-Oriented Database Systems x86 is a high-level language Introduction to HAMT Sirius: An Open End-to-End Voice and Vision Personal Assistant and Its Implications for Future Warehouse Scale Computers Staring into the Abyss: An Evaluation of Concurrency Control with One Thousand Cores Paper: Staring into the Abyss: An Evaluation of Concurrency Control with One Thousand Cores Compressed Text Indexes: From Theory to Practice! Command-line tools can be 235x faster than your Hadoop cluster Predecessor search for Big Data: x-fast tries, locality of reference and all that Understanding and Expressing scalable Concurrency The Art of Approximating Distributions: Histograms and Quantiles at Scale q-digest - Medians and Beyond: New Aggregation Techniques for Sensor Networks q-digest : an algorithm for computing approximate quantiles on a collection of integers References for Data Stream Algorithms A closer Look at GPUs Lecture 14 - Out-of-Order Execution - Carnegie Mellon - Computer Architecture 2013 - Onur Mutlu Operation fusion and deforestation for Scala Specialized Evolution of the General-Purpose CPU Immutability Changes Everything The Design of Approximation Algorithms Scalability! But at what COST? Programming on Parallel Machines - GPU, Multicore, Clusters and More Mergeable persistent data structures High Performance Hardware-Accelerated Flash Key-Value Store SipHash: a fast short-input PRF Analysis of Pivot Sampling in Dual-Pivot Quicksort, A Holistic Analysis of Yaroslavskiy’s Partitioning Scheme Suffix Trees and their Applications in String Algorithms Borislav Petkov: x86 instruction encoding and the nasty hacks we do in the kernel Don’t Thrash: How to Cache Your Hash on Flash Resizable, Scalable, Concurrent Hash Tables via Relativistic Programming Cuckoo Filter: Practically Better than Bloom Resizable, Scalable, Concurrent Hash Tables via Relativistic Programming What's the deal with Hardware Transactional Memory!?! (linux.conf.au 2014) SimHash or the way to compare quickly two datasets An Overview of Kernel Lock Improvements (on huge NUMA systems) Invertible Bloom Lookup Tables Instruction Sets Should Be Free: The Case For RISC-V Accurate Methods for the Statistics of Surprise and Coincidence Epiphany Architecture Reference Perceptual Hashing (searching similar images, reverse image search) The Bw-Tree: A B-tree for New Hardware Platforms Philip Wadler: Why no one uses functional languages Rank and select for succinct data structures Bridging Islands of Specialized Code using Macros and Reified Types Data types ala carte Throw away the keys: Easy, Minimal Perfect Hashing MICA: A Holistic Approach to Fast In-Memory Key-Value Storage SQL versus coSQL — a compendium to Erik Meijer’s paper Mark Hill CPU, TLB Similarity Measurement on Leaf-labelled Trees Modern Microprocessors - A 90 Minute Guide! Is Parallel Programming Hard, And, If So, What Can You Do About It?, Hardware and its Habits Why Functional Programming Matters Fast Computation of min-Hash Signatures for Image Collections Scalable, Example-Based Refactorings with Refaster Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture AA-Sort: A New Parallel Sorting Algorithm for Multi-Core SIMD Processors An Experimental Study of Sorting and Branch Prediction Data Structures and Algorithms for Nearest Neighbor Search in General Metric Spaces (VP-tree) Duplicate News Story Detection Revisited Monoids: Theme and Variations (Functional Pearl) Amazon.com Recommendations - Item-to-Item Collaborative Filtering HyperLogLog: the analysis of a near-optimal cardinality estimation algorithm Phil Bagwell, Ideal Hash Trees, 2001 Phil Bagwell, Fast And Space Efficient Trie Searches, 2000 HyperLogLog in Practice: Algorithmic Engineering of a State of The Art Cardinality Estimation Algorithm Memory system compression and its benefits The Night Watch Miniboxing: Improving the Speed to Code Size Tradeoff in Parametric Polymorphism Translations Iterative Ranking from Pair-wise Comparisons Fast Mergable Integer Maps Sketch of the Day: Frugal Streaming (median, rank) TRASH A dynamic LC-trie and hash data structure RAY: Integrating Rx and Async for Direct-Style Reactive Streams Instruction tables Lists of instruction latencies, throughputs and micro-operation break-downs for Intel, AMD and VIA CPUs Memory Barriers: a Hardware View for Software Hackers Monads for functional programming RRB-Trees: Efficient Immutable Vectors Five Myths about Hash Tables dablooms - an open source, scalable, counting bloom filter library