Clash of the Titans: MapReduce vs. Spark for Large Scale Data Analytics

Gorilla: A Fast, Scalable, In-Memory Time Series Database

TreeScope: Finding Structural Anomalies in Semi-Structured Data

KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing

Knowledge-Based Trust: Estimating the Trustworthiness of Web Sources

One trillion edges: graph processing at Facebook scale

Understanding the Causes of Consistency Anomalies in Apache Cassandra [pdf]

An Evaluation of Concurrency Control with One Thousand Cores [pdf]

Minuet: A Scalable Distributed Multiversion B-Tree [pdf]

Coordination Avoidance in Distributed Systems