Intelligence artificielle et données

307 documents à télécharger gratuitement

Cours, examens, TD, TP et exercices de intelligence artificielle et données. Thèmes couverts : machine learning, apprentissage automatique, big data, data science, fouille de données.

M1 BADS

This document presents a multidimensional analysis of food data and socio-professional categories. It includes correlation matrices and results from Principal Component Analysis (PCA). The aim is to study the relationships between different food categories and socio-professional groups.

autre
ordinaire
pain
6p0
Apache Pig: A High-Level Platform for Large-Scale Data Analysis

Apache Pig is a high-level platform for creating MapReduce programs used with Hadoop, utilizing a procedural language called Pig Latin. Pig Latin abstracts Java-based MapReduce programming into a more flexible and SQL-like syntax with control over execution flow and data transformations. It supports parallel processing, user-defined functions (UDFs) in multiple languages, and handles large-scale datasets efficiently while combining aspects of both declarative SQL and procedural MapReduce. Advantages include a quick learning curve, flexibility in data processing, and the ability to embed cus...

Pig Latin
MapReduce
SQL
62p0
MapReduce

This document introduces the MapReduce paradigm, a distributed computing model to process large datasets, especially in scenarios involving tasks such as data aggregation and analysis. Traditional sequential methods are contrasted with the efficiency and parallelization enabled by MapReduce, with emphasis on how large-scale problems can be divided into smaller subtasks. The two core operations, MAP and REDUCE, are described in depth, along with examples in word frequency counting, web analytics, and finding mutual connections in social graphs. The advantage of automating data distribution a...

MapReduce
Hadoop
Divide and Conquer
88p0
MapReduce in Big Data

This document explores the MapReduce paradigm, designed for processing large data sets through distributed computing. It discusses the traditional sales aggregation problem, outlining inefficiencies and presenting MapReduce as a solution. The framework allows for efficient problem-solving by breaking down tasks and parallelizing the computation necessary for big data analysis.

quot
donn
data
88p0
Chapitre III : MapReduce

This document provides an in-depth explanation of the MapReduce paradigm used in large-scale distributed computing. It describes the inefficiencies of traditional solutions to analyze large datasets and introduces MapReduce as an effective alternative. The methodology involves dividing tasks into smaller, parallelizable operations: `MAP` to transform data into key-value pairs and `REDUCE` to aggregate results by key. Various examples, including word frequency analysis, website statistics, and common connections in social graphs, illustrate practical applications of the MapReduce model. The...

MapReduce
Hadoop
Distributed Computing
88p0
MapReduce

This document provides an overview of the MapReduce programming model and its efficiency in processing large datasets. It discusses the traditional methods of data processing and introduces the MapReduce paradigm, emphasizing its divide and conquer approach. The document outlines the operations of MAP and REDUCE, detailing the steps involved in the MapReduce process.

quot
donn
data
88p0
Hadoop: An Overview of Big Data Framework

Hadoop is an open-source framework designed for distributed data processing across massive datasets, with scalable and fault-tolerant architecture. Core components include HDFS for distributed storage and MapReduce for distributed computing, with additional tools like Hive, Pig, and HBase for advanced analytics and data management. Inspired by Google's publications, Hadoop leverages principles like data redundancy, clustering, and parallel processing to support Big Data challenges. It has been widely adopted across industries, including by organizations such as Facebook, Google, and Amazon.

Hadoop
MapReduce
HDFS
42p0
Big Data Course - Chapter 2: Hadoop

Hadoop, an open-source framework managed by Apache and written in Java, facilitates distributed processing of massive datasets using clusters of commodity hardware. Key components include HDFS for distributed storage and MapReduce for parallelized data processing, ensuring fault tolerance, scalability, and cost-efficiency. The ecosystem extends functionality through tools such as Pig, Hive, and Flume for data processing, storage, scheduling, and monitoring. Developed by Doug Cutting and inspired by Google technologies, Hadoop powers numerous enterprises like Amazon, Adobe, Facebook, and Twi...

Hadoop
MapReduce
HDFS
42p0
Hadoop

Hadoop is an open-source framework designed for distributed processing of large data sets across clusters. It facilitates the creation of applications that handle massive volumes of data while ensuring performance and fault tolerance. The technology has evolved under the Apache foundation and is widely used in various enterprises.

hadoop
hdfs
namenode
42p0
Apache Spark: Framework, Advantages, and Ecosystem

This document provides an in-depth critique of MapReduce's limitations, emphasizing its inefficiencies in complex, multi-step computations. It introduces Apache Spark as a more advanced, memory-optimized solution with higher performance, scalability, and interactive APIs for diverse programming languages. Key Spark components like RDDs, DataFrames, Datasets, and its ecosystem libraries (e.g., MLlib, GraphX) are detailed alongside their integration capabilities. The document further outlines Spark's architecture, including its execution model and resource management, while contrasting its pe...

Apache Spark
MapReduce
Big Data
47p0
Comparaison MapReduce vs Spark pour Big Data

This document discusses the limitations of MapReduce for complex tasks and introduces Apache Spark as a more efficient alternative. It covers the features, ecosystem, and architecture of Spark, emphasizing its ability to handle both batch and real-time processing. Spark is designed for speed, utilizing in-memory data storage, and supports various programming languages for ease of application development.

spark
donn
quot
47p0
Techniques of Incremental ETL

The document explores various techniques for performing incremental ETL (Extract, Transform, Load) processes. It outlines three architectural approaches: single-stage ETL for small-scale BI solutions with minimal data sources, two-stage ETL for moderately complex or larger data volumes that a single stage cannot accommodate, and three-stage ETL for even more complex architectures that optimize data flow and reduce load on source systems. The content is designed to guide data engineers in building efficient ETL pipelines tailored to specific use cases.

ETL
incremental ETL
data architecture
4p0
Incremental ETL Techniques Explained

The document explores incremental ETL techniques with three specific architectures: one-stage, two-stage, and three-stage ETL processes. A one-stage ETL is suitable for small-scale BI systems with limited data sources and simple requirements. Two-stage ETL addresses complexities and larger data volumes, optimizing performance. Lastly, a three-stage ETL architecture reduces extraction overhead and minimizes source system load, enhancing efficiency in data flow management.

Incremental ETL
ETL stages
BI architecture
4p0
Introduction aux Big Data

This document provides an overview of Big Data, including its definition, sources, challenges, and the 5 Vs that characterize it. It highlights the importance of data in decision-making processes and the necessity of effectively managing and analyzing this data. Additionally, it discusses the rapid generation of large volumes of unstructured data and the implications for businesses and technology.

donne
data
introduction
17p0
Informatique Décisionnelle : Le data Warehouse

G n ralit s sur les Data Warehouses Introduction Dans le monde des affaires actuel, les donn es et l'analyse jouent un r le indispensable dans le processus de prise de d cision.

Informatique
exam
est
1p0
TP9: Utilisation des expressions régulières en R pour le traitement des données email

This document explores the application of regular expressions in R for processing email data. Initially, challenges in parsing email domains are identified, leading to the utilization of specific regular expressions to extract data systematically. Functions like 'grep', 'gsub', 'regexpr', and 'regmatches' are discussed and applied to manipulate, search, and modify email strings effectively. Results include generating a frequency table for email domains from the processed dataset.

regular expressions
R programming
data parsing
2p0
Data Processing with Regular Expressions in R

This document discusses data reading, the application of regular expressions in R, and how to extract domains from email addresses. It provides coding examples and explanations of the syntax used in regular expressions. The focus is on overcoming challenges associated with unique email formats and counting domain frequencies.

quot
email
nous
2p0
Introduction to Data Import and Manipulation in R

The document provides an overview of methods to import, read, and manipulate tabular data formats in R, such as CSV and Excel files. It discusses the use of built-in functions like `read.csv` and libraries like `readxl` for handling external datasets. Methods to access specific rows, columns, and subsets of these data frames are thoroughly explained. Additionally, the document introduces built-in datasets in R and demonstrates their usage with key functions such as `data()` and `help()`.

R programming
CSV
Excel files
3p0
Apprentissage avec Python et Scikit-learn

This document explains ensemble learning techniques (bagging, random forests, and boosting) for classification and regression problems using Python and Scikit-learn. Bagging involves averaging predictions of multiple predictors to reduce variance and works well with strong predictors. Random forests enhance decision trees by introducing randomness in attribute selection, further reducing variance. Boosting sequentially trains weak classifiers, modifying data distributions to obtain strong combined results. The lab exercises include variance analysis on model accuracy, hyperparameter tuning,...

ensemble learning
BaggingClassifier
boosting
4p0
Apprentissage avec Python et Scikit-Learn

This document focuses on decision trees for regression using Scikit-learn, guiding the user through creating a regressor and its application. A sinusoidal signal with added noise is used as a dataset to train and test the regression model. Key experiments include analyzing the impact of the decision tree's maximum depth and the level of noise in the dataset, along with an exploration of the Diabetes dataset. The report also involves evaluating model performance using mean squared error and implementing a grid search for hyperparameter tuning.

Decision trees
DecisionTreeRegressor
Regression
4p0
MODULE 3 : Intégration de données

This document provides an in-depth overview of data integration techniques in Business Intelligence (BI) systems. It explores the ETL (Extract, Transform, Load) process, EII (Enterprise Information Integration), and EAI (Enterprise Application Integration), detailing their characteristics, advantages, disadvantages, and use cases. Core methodologies include data extraction (both real-time and batch modes), data transformations, and strategies for resolving data inconsistencies. Additionally, examples of commercial tools, such as Oracle Warehouse Builder and IBM WebSphere, are highlighted fo...

ETL (Extract
Transform
Load)
58p0
Module 3: Intégration de données

Ce module traite de l'intégration des données dans le contexte de la Business Intelligence. Il aborde les processus ETL, les outils d'intégration et les différences entre ELT et ETL. Les défis liés à l'intégration de données provenant de sources disparates sont également discutés.

donn
tarbout
sources
58p0
TP : Génération des outputs sur SPSS

This document describes a series of exercises focused on generating and interpreting outputs using SPSS statistical software. Tasks include creating frequency tables and cross-tabulations for variables such as travel purpose, global satisfaction, and respondent demographics (e.g., gender, profession). Graphs such as histograms and pie charts are created to visualize data, and analytical insights involve calculating age statistics and average scores for key criteria (ticket price, service quality, and flight punctuality). The exercises emphasize cross-tabulation with percentages for deeper i...

SPSS
frequency tables
cross-tabulation
2p0
Devoir Surveillé - Système d’Information Décisionnel

The document discusses a decision-support system module with a focus on data warehouse design for the Tunisian Postal Service. The theoretical portion includes comparisons of data warehouse architectures and scenarios where implementing a data warehouse is unsuitable. A case study emphasizes designing a data warehouse for payment tracking via smart payment cards. The proposed data model needs to support comprehensive reporting on financial metrics such as revenue by provider or zone, client segmentation by transaction behaviors, and promotional campaign targeting.

Data warehouse
R. Kimball architecture
B. Inmon architecture
2p0

Autres ressources en intelligence artificielle et données