What is Big Data

Опубликовано: 28 Октябрь 2024
на канале: Technically U
29
2

Big Data is capturing, storing, processing, analyzing, and extracting a large amount of information that continues to grow exponentially over time. This information can come from several sources including the Internet of Things (IoT), sensors, smart devices (watches, phones, cameras, tablets), social media, traffic devices, surveillance videos, utility meters, point-of-sale terminals, photos, music, health records, web searches and this source of information continues to grow.
Businesses want and need this information since it can be used to provide valuable insights and a competitive edge, detect fraudulent activities, gain operational efficiencies, improve decision-making processes, provide better customer services, and help them rapidly deploy new applications and products.
Most standard data processing software and systems cannot scale to process efficiently the large and complex information captured by Big data sets, nor can the traditional IT staff members manage it. Big Data requires extraordinarily complex machinery and data scientist to extract meaningful and consumable information from all this collected data.
The Characteristics of Big Data
Volume
The volume of data plays a role in determining if valuable information can be pulled from it. The size of Big Data ranges from terabytes, petabytes, to zettabytes. As mentioned earlier this data can be collected from a variety of sources.
Variety
Variety refers to the many types of data that are available. Structured, Semi-Structured, and Unstructured. Data from emails, photos, videos, IoT devices may be gathered as raw unstructured data requiring special processing and analysis to make it meaningful and useful.
Velocity
Velocity refers to the speed at the data is generated. The massive amount of continuous data generated needs to be collected, stored, process analyzed, and acted on as quickly as possible while it is still valuable to the organization
Variability
Variability refers to the inconsistency and unpredictability of data flows. The amount of data flows varies depending on the time of day, captured activities, social media trends, and seasonal activities.
Veracity
Veracity refers to the quality and accuracy of the collected data. If source data is incorrect or corrupted, analyses will be worthless.

There are Three Types of Big Data
Structured
Structured data is known as quantitative data. This type of data can be stored, organized, accessed, and processed in a fixed format. Structured data is valuable because you can gain insights into overarching trends by running the data through data analysis methods, such as regression analysis and pivot tables.
Unstructured
Unstructured data or qualitative data is any data with an unknown form. Unstructured data poses multiple challenges when it comes to processing it for valuable and meaningful data.
Semi-structured
Semi-structured data is information that is not sourced from a relational database or any other data table but contains organizational properties that make it easy to analyze.
In Conclusion
Big Data is capturing, storing, processing, analyzing, and extracting a large amount of information that continues to grow exponentially over time. This information can come from several sources. The characteristics of Big Data are Volume, Variety, Velocity, Variability, and Veracity. Big Data comes in three different fashions: Structured, Semi-structured, and Unstructured. Some major advantages of Big Data are gaining a competitive edge, detecting fraudulent activities, gaining operational efficiencies, improving decision-making processes, better customer services, and rapidly deploying new applications and products.