GRAPH NAME DATA AND DISTRIBUTED SYSTEMS *** ## NODE 1 NAME IBM 305 RAMAC DATE 1956-09-14 PLACE USA - California - San Jose WHO IBM BRIEF DESCRIPTION IBM introduces the 305 RAMAC, one of the earliest commercial systems to use magnetic disk storage with direct access. The technology changes data processing by allowing records to be accessed without depending exclusively on sequential media. LINK *** ## NODE 2 NAME INTEGRATED DATA STORE DATE 1963 PLACE USA WHO Charles Bachman BRIEF DESCRIPTION Charles Bachman develops Integrated Data Store, one of the earliest general-purpose database management systems. Its navigational model based on records and relationships later influences the network database model. LINK *** ## NODE 3 NAME IBM IMS DATE 1968 PLACE USA WHO IBM BRIEF DESCRIPTION IBM develops the Information Management System as a hierarchical database management system. IMS enables the management of large volumes of structured records and represents an important stage in the evolution of enterprise data systems. LINK *** ## NODE 4 NAME RELATIONAL MODEL DATE 1970 PLACE USA - California - San Jose WHO Edgar F. Codd BRIEF DESCRIPTION Edgar F. Codd proposes representing data through mathematical relations organized as sets of tuples. The model separates the logical representation of data from physical storage details and establishes the conceptual foundation of relational databases. LINK *** ## NODE 5 NAME SYSTEM R DATE 1974 PLACE USA - California - San Jose WHO IBM BRIEF DESCRIPTION IBM begins System R as an experimental project to demonstrate that the relational model can be implemented efficiently. The system contributes to the development of query optimization, transaction processing, and declarative data-access languages. LINK *** ## NODE 6 NAME SEQUEL DATE 1974 PLACE USA - California - San Jose WHO Donald Chamberlin and Raymond Boyce BRIEF DESCRIPTION SEQUEL is developed as a declarative language for querying relational databases in System R. It later evolves into SQL, which becomes a standard interface for defining, querying, and manipulating relational data. LINK *** ## NODE 7 NAME ENTITY-RELATIONSHIP MODEL DATE 1976 PLACE USA WHO Peter Chen BRIEF DESCRIPTION Peter Chen formalizes the entity-relationship model for conceptually representing entities, attributes, and relationships before implementing a database. The approach becomes a fundamental tool for modeling and designing information systems. LINK *** ## NODE 8 NAME LOGICAL CLOCKS DATE 1978 PLACE USA WHO Leslie Lamport BRIEF DESCRIPTION Leslie Lamport formalizes a precedence relationship between events in distributed systems and proposes logical clocks for partially ordering events without depending on a global physical clock. The work establishes foundations for reasoning about causality and distributed coordination. LINK *** ## NODE 9 NAME COMMERCIAL SQL DATABASE DATE 1979 PLACE USA - California WHO RELATIONAL SOFTWARE INC. BRIEF DESCRIPTION Relational Software Inc., later Oracle Corporation, commercializes a relational database implementation compatible with SQL. Commercial adoption helps extend the relational model beyond experimental research environments. LINK *** ## NODE 10 NAME BYZANTINE GENERALS DATE 1982 PLACE USA WHO Leslie Lamport, Robert Shostak and Marshall Pease BRIEF DESCRIPTION The Byzantine Generals Problem formalizes the difficulty of reaching agreement when some components of a distributed system may produce incorrect or contradictory information. The model becomes a fundamental reference for distributed fault tolerance. LINK *** ## NODE 11 NAME FLP CONSENSUS IMPOSSIBILITY DATE 1985 PLACE USA WHO Michael Fischer, Nancy Lynch and Michael Paterson BRIEF DESCRIPTION The FLP result demonstrates that in a completely asynchronous distributed system, no deterministic algorithm can guarantee consensus if even a single process may fail by stopping. The result establishes a fundamental limit for distributed protocols. LINK *** ## NODE 12 NAME SQL STANDARD DATE 1986 PLACE USA WHO ANSI BRIEF DESCRIPTION SQL receives its first formal standardization as a language for relational databases. Standardization facilitates application portability and establishes a common interface among database management systems developed by different vendors. LINK *** ## NODE 13 NAME CORBA DATE 1991 PLACE USA WHO OBJECT MANAGEMENT GROUP BRIEF DESCRIPTION CORBA defines a middleware architecture that enables distributed objects written in different languages and running on heterogeneous systems to invoke operations on one another through standardized interfaces. LINK *** ## NODE 14 NAME PAXOS DATE 1998 PLACE USA WHO Leslie Lamport BRIEF DESCRIPTION Paxos formalizes a family of algorithms for achieving consensus among distributed processes even when failures may occur. Its principles directly influence replicated systems, coordination services, and distributed databases. LINK *** ## NODE 15 NAME CAP DATE 2000 PLACE USA - California WHO Eric Brewer BRIEF DESCRIPTION Eric Brewer proposes the relationship among consistency, availability, and partition tolerance in distributed data systems. The result, later formalized, establishes fundamental constraints on the properties that can be guaranteed simultaneously during network partitions. LINK *** ## NODE 16 NAME GOOGLE FILE SYSTEM DATE 2003 PLACE USA - California WHO GOOGLE BRIEF DESCRIPTION Google describes a distributed file system designed to store large datasets across numerous commodity servers. The architecture uses replication, block distribution, and failure recovery as normal operational properties. LINK *** ## NODE 17 NAME MAPREDUCE DATE 2004 PLACE USA - California WHO GOOGLE BRIEF DESCRIPTION Google presents MapReduce as a programming model for parallel processing of large volumes of distributed data. The model abstracts distribution, coordination, and failure recovery through conceptual map and reduce phases. LINK *** ## NODE 18 NAME AMAZON S3 DATE 2006-03-14 PLACE USA WHO AMAZON WEB SERVICES BRIEF DESCRIPTION Amazon S3 introduces object storage accessible as a service through network interfaces. The model separates storage from physical infrastructure directly managed by the user and becomes an important component of cloud architectures. LINK *** ## NODE 19 NAME APACHE HADOOP DATE 2006 PLACE USA WHO APACHE SOFTWARE FOUNDATION BRIEF DESCRIPTION Hadoop emerges as an open platform for distributed storage and processing of large datasets. Its architecture incorporates a distributed file system and an execution model initially inspired by MapReduce. LINK *** ## NODE 20 NAME GOOGLE BIGTABLE DATE 2006 PLACE USA - California WHO GOOGLE BRIEF DESCRIPTION Google presents Bigtable, a distributed system for storing structured data at large scale through a sparse, multidimensional table model. Its architecture influences numerous later distributed databases. LINK *** ## NODE 21 NAME AMAZON EC2 DATE 2006-08-25 PLACE USA WHO AMAZON WEB SERVICES BRIEF DESCRIPTION Amazon EC2 provides virtualized computing capacity on demand through a service interface. The model allows computing resources to be provisioned and released dynamically and represents a milestone in the consolidation of infrastructure as a service. LINK *** ## NODE 22 NAME AMAZON DYNAMO DATE 2007 PLACE USA WHO AMAZON BRIEF DESCRIPTION Amazon describes Dynamo, a distributed key-value store designed to maintain high availability through partitioning, replication, and eventual consistency. Its architecture influences the later development of distributed NoSQL systems. LINK *** ## NODE 23 NAME APACHE CASSANDRA DATE 2008 PLACE USA - California WHO FACEBOOK BRIEF DESCRIPTION Cassandra is initially developed at Facebook as a distributed storage system designed for large volumes of data. It combines partitioning and replication techniques inspired by systems such as Dynamo and Bigtable. LINK *** ## NODE 24 NAME APACHE KAFKA DATE 2011 PLACE USA - California WHO LINKEDIN BRIEF DESCRIPTION Kafka is developed to transport and retain large-scale event streams through partitioned and replicated distributed logs. Its model enables producers and consumers to remain decoupled and provides infrastructure for data integration and event processing. LINK *** ## NODE 25 NAME GOOGLE SPANNER DATE 2012 PLACE USA - California WHO GOOGLE BRIEF DESCRIPTION Google describes Spanner as a globally distributed database combining partitioning, replication, and temporal coordination to provide transactions and external consistency at large scale. LINK *** ## NODE 26 NAME APACHE SPARK DATE 2014 PLACE USA - California WHO APACHE SOFTWARE FOUNDATION BRIEF DESCRIPTION Apache Spark becomes established as a distributed data-processing platform using abstractions that allow working datasets to remain in memory and iterative analytical workloads to execute with less dependence on intermediate disk storage. LINK *** ## NODE 27 NAME KUBERNETES DATE 2014-06-06 PLACE USA - California WHO GOOGLE BRIEF DESCRIPTION Google releases Kubernetes as an open-source platform for deploying, coordinating, and managing containerized applications across clusters of machines. The system automates scheduling, replication, recovery, and updates of distributed services. LINK