Course Purpose

This course aims to equip learners with essential knowledge and skills to manage, process, and analyze large-scale data for informed decision-making. It introduces key Big Data concepts, including data types, characteristics, and analytics processes, while emphasizing their practical application in real-world contexts.

The course also provides hands-on experience with tools in the Apache Hadoop ecosystem, enabling learners to work with distributed data systems. By the end of the course, learners will be able to apply Big Data techniques to solve problems and support organizational performance.


 

 

Course Learning Outcomes

 Describe key Big Data concepts, including the 5Vs and system architectures. 

Use Big Data tools (e.g., Hadoop, Spark) to perform data processing tasks. 

Interpret and analyze datasets to identify trends and support decision-making. 

Develop a basic Big Data analytics project or solution.

 

Course Content

Big data analytics: Definitions, terminology of Big data, characteristics of Big data, Big data overview and trends, drivers of Big data, types of Big data technology, and Big data analytics lifecycle

Big data technologies: Data storage, data mining, data analytics, data visualization, and emerg- ing Big data technologies

Big data analysis: Handling and processing Big data, Methodological challenges and prob- lems, Big data analytics and machine learning, analysis and applications

Hadoop: Definitions, history, hadoop main components, key characteristics of Hadoop, cur- rent status, and Leverage Hadoop Distributed File System (HDFS): Goals, Hadoop server roles, Hadoop cluster, Spark technologies, data modeling and storage, HDFS details, and HDFS user interface

Hadoop Map-Reduce: Terminologies, MapReduce programming, Hadoop streaming, and MapRe- duce and HDFS