Computing and Multimedia Technology

Large-Scale Data Storage

<< back to Curriculum Plan

6 ECTS; 3º Ano, 1º Semestre, 28,0 PL + 28,0 TP + 5,0 OT , Cód. 814360.

Lecturer
- Renato Eduardo Silva Panda (1)(2)

(1) Lead Professor
(2) Teaching Professor

Prerequisites
Not applicable.

Objectives
This course develops skills in data acquisition, preparation, storage, querying and analysis, connecting data science with data engineering and large-scale processing. By the end of the course, students should be able to:

1. Explain the data project lifecycle, the characteristics of Big Data and the ethical implications of collecting and using data.
2. Configure reproducible development environments and use Python and its libraries to manipulate and analyse data.
3. Acquire and integrate data from files, APIs and web pages, applying preparation and transformation operations with data quality checks.
4. Conduct exploratory analysis and communicate findings through visualisations, reports and interactive dashboards.
5. Compare storage and querying models, selecting formats and technologies suited to data characteristics and usage requirements.
6. Explain parallel and distributed processing principles and apply large-scale processing strategies, assessing their costs and limitations.
7. Develop, document and defend practical data solutions, justify technical decisions, and critically verify results and assistance obtained from artificial intelligence tools.

Program
1. Introduction to data science and Big Data
- 1.1 Fundamental concepts, professional roles and the data project lifecycle.
- 1.2 The 5Vs: volume, velocity, variety, veracity and value.
- 1.3 Ethics, privacy, transparency, data quality and social impact.
- 1.4 Reproducibility, documentation, version control and critical use of AI.
2. Development environment and Python fundamentals
- 2.1 Jupyter, Python scripts and Visual Studio Code.
- 2.2 Isolated environments and dependency management: pip, conda and uv.
- 2.3 Docker, Docker Compose and, if time allows, an introduction to Dev Containers.
- 2.4 Python review: types, data structures, control flow, functions and modules.
- 2.5 Project organisation, Git, code checks and basic tests.
3. Data manipulation, analysis and visualisation
- 3.1 NumPy and array operations.
- 3.2 pandas: selection, transformation, aggregation and joins of tabular data.
- 3.3 Exploratory analysis, descriptive statistics, missing data and outliers.
- 3.4 Matplotlib, Seaborn and Plotly; chart selection and accurate communication.
- 3.5 Interactive dashboards with Streamlit and an introduction to alternatives such as Dash.
4. Data acquisition, integration and engineering
- 4.1 Files and formats: CSV, JSON, Excel and Parquet.
- 4.2 REST APIs: authentication, pagination, rate limits and error handling.
- 4.3 Web scraping with requests and Beautiful Soup; good collection practices.
- 4.4 Document data extraction and an introduction to OCR.
- 4.5 Cleaning, normalisation, integration and validation of data from multiple sources.
- 4.6 Introduction to ETL and ELT, traceability and data updates, with guided examples.
5. Data storage and querying
- 5.1 Relational and NoSQL models: documents, key-value, column families and graphs.
- 5.2 SQL querying and analysis; examples with PostgreSQL and DuckDB.
- 5.3 Storage and querying with MongoDB and Redis.
- 5.4 Object storage, data warehouses, data lakes and an introduction to lakehouses.
- 5.5 Selection criteria: data model, access patterns, consistency, scalability and cost.
- 5.6 Comparative experiments with notebooks and local Docker services, where applicable.
6. Large-scale processing strategies
- 6.1 Memory and I/O constraints; assessing execution time and resource usage.
- 6.2 Chunking, vectorised operations and incremental processing.
- 6.3 Columnar formats, compression, partitioning and selective reads.
- 6.4 Parallelisation, lazy execution and an introduction to Dask.
7. Distributed processing
- 7.1 MapReduce and the fundamentals of distributing data and tasks.
- 7.2 Hadoop architecture: HDFS and YARN; the transition to Spark.
- 7.3 Apache Spark and PySpark: DataFrames, Spark SQL and an introduction to RDDs.
- 7.4 Transformations, actions, lazy execution, partitions and data transfers between tasks.
- 7.5 Comparing local, parallel and distributed solutions; measuring and interpreting performance.
- 7.6 Streaming and an introduction to distributed machine learning with Spark MLlib, if time allows.

Evaluation Methodology
All assessment periods:

- 30%: Project I, Data Science: acquisition, preparation and exploratory analysis (6 points).
- 30%: Project II, Big Data: storage, querying and large-scale processing (6 points).
- 40%: Theoretical examination (8 points).

Minimum grades:

- 50% in each project: 10/20, equivalent to 3/6.
- 35% in the examination: 7/20, equivalent to 2.8/8.
- A final grade of at least 9.5/20, with all component minima met.

For grades out of 20, the final grade is 0.30 × P1 + 0.30 × P2 + 0.40 × E. Contributions graded out of 6, 6 and 8 are added together.

Both projects must be submitted and defended by the deadlines and at the times set during continuous assessment. Projects cannot be repeated or replaced by an examination. Failure to submit or defend any project excludes the student from course assessment in all periods of that academic year. Failure to meet any project minimum prevents a passing grade.

Project grades carry over across assessment periods within the same academic year. Only the theoretical examination may be repeated in applicable periods. Defences verify individual understanding and contributions, including AI use.

Bibliography
- McKinney, W. (2022). Python for Data Analysis: Data Wrangling with pandas, NumPy, and Jupyter. .: O'Reilly Media
- Triguero, I. e Galar, M. (2023). Large-Scale Data Analytics with Python and Spark. UK: Cambridge University Press

Teaching Method
TP classes combine concepts and practical tutorials. PL classes provide self-paced notebook and Docker experiments and support for two projects. Independent work and critical AI use support explanations, learning and solution verification.

Software used in class
Python, Jupyter, Visual Studio Code, pip, conda/Miniconda or micromamba, uv, Git, Docker and Docker Compose; NumPy, pandas, Matplotlib, Seaborn, Plotly, Streamlit, requests and Beautiful Soup; SQL, DuckDB, PostgreSQL, MongoDB, Redis, Dask and PySpark. Ruff and pytest support code quality and verification. Dev Containers, Dash and document extraction/OCR tools are used in supplementary examples.

 

 

 


<< back to Curriculum Plan
ISO 9001
NP4552
SGC
KreativEu
erasmus
catedra
b-on
portugal2020
centro2020
compete2020
crusoe
fct
feder
fse
poch
portugal2030
poseur
prr
santander
republica
UE next generation
Centro 2030
Lisboa 2020
Compete 2030
co-financiado