Book Profile
Big Data_ A Very Short Introduction (Very Short Introductions)
Very Short Introductions
A concise introduction to what big data is, how it is collected, stored, and analysed, and how it is transforming medicine, business, security, and society.
Get the book →Big Data: A Very Short Introduction demystifies one of the defining technological forces of our age, charting how data evolved from notched bones and census tallies to the exabyte-scale streams of the digital universe. Dawn Holmes explains, in plain language and with clear diagrams, what distinguishes big data from traditional 'small data'—volume, variety, velocity, and veracity—and how new storage architectures (Hadoop, NoSQL, the Cloud) and analytic techniques (clustering, classification, MapReduce, Bloom filters, PageRank, recommender systems) extract useful information from massive, messy datasets. Through vivid case studies—Google Flu Trends, the Ebola and Nepal disaster responses, Amazon, Netflix, the Snowden leaks, WikiLeaks, and the rise of smart homes and cities—the book shows both the immense promise and the real perils of a data-driven world, ending with a call to use big data's power responsibly.
What it argues
Big Data_ A Very Short Introduction (Very Short Introductions)
Key ideas it contributes
- Data Volume, Velocity, and Variety — The scale, speed of generation, and heterogeneity (structured, semi-structured, unstructured) of data produced across search engines, sensors, social media, and other digital sources that together characterize big data.
- Data Veracity (Quality) — The accuracy, reliability, and trustworthiness of collected data, recognizing that digital-age data is often imprecise, uncertain, biased, or simply untrue and requires pre-processing for consistency.
- Distributed Storage Architecture — The design choice of scalable storage and management systems such as Hadoop distributed file systems, NoSQL databases, and Cloud infrastructure that enable horizontal scalability and fault tolerance for big data.
- Analytic Technique Application — The application of data mining and machine learning methods such as clustering, classification, MapReduce, Bloom filters, PageRank, and recommender algorithms to discover patterns and extract knowledge from big data.
- Useful Information Extraction — The intermediate state in which raw, often unstructured data is transformed into meaningful, valuable information and patterns that can inform understanding, prediction, and action.
- Predictive Model Accuracy — The degree to which models built from big data correctly forecast outcomes, sensitive to model construction issues such as over-fitting, spurious correlation, and failure to update for changing conditions.
- Data Security and Privacy Measures — The protective controls such as encryption, firewalls, access authentication, and anonymization deployed to safeguard data from theft, tampering, hacking, and unauthorized disclosure.
- Decision Quality and Organizational Value — The downstream benefits of big data including improved business decisions, targeted marketing, better patient care, cost reduction, scientific discovery, and societal efficiency gains in smart homes and cities.
- Privacy and Security Risk — The adverse outcome of exposure to data theft, breaches, surveillance, identity theft, and loss of personal privacy arising from large-scale collection and storage of personal data.