Travaux pratiques (TP) - Intelligence artificielle et données

22 documents à télécharger gratuitement

Travaux pratiques (TP) de intelligence artificielle et données, partagés par des étudiants et des enseignants. Thèmes couverts : machine learning, apprentissage automatique, big data, data science, fouille de données.

Atelier Fouille de Données: Les Méthodes de Classification Non Supervisées

This document explores unsupervised classification methods, focusing primarily on two types: hierarchical and non-hierarchical methods. It details the K-means algorithm, explaining its principles, convergence criteria, and the role of intra-class and inter-class variance. Hierarchical clustering is also discussed, describing dendrograms and aggregation criteria such as single linkage, complete linkage, and variance-based methods. Additionally, the document provides an overview of mixed classification, combining K-means partitioning with hierarchical clustering for a more refined analysis.

unsupervised classification
K-means algorithm
hierarchical clustering
33p0
Atelier Fouille de données

This document outlines methods of unsupervised classification, focusing on automatic classification techniques. It discusses the concepts of dissimilarity and inertia as they relate to grouping individuals based on their characteristics. Additionally, it presents the K-means algorithm for classifying data.

classification
rarchique
classes
33p0
Mastère Business Analytics & Data Science - Atelier Fouille de données

This document outlines a course workshop on data mining as part of the Business Analytics & Data Science master's program. It includes a section on principal component analysis (PCA) applied to vehicle data with specific tasks for analysis. The focus is on interpreting correlations and distributions of vehicles using PCA techniques.

analyse
composantes
principales
8p0
Analyse en Composantes Principales (ACP) Atelier

The document outlines a practical workshop for Principal Component Analysis (PCA) applied to six continuous variables recorded from data on 20 cars. Using R programming and the FactoMineR and factoextra packages, correlations between variables are analyzed to identify patterns and influences in the factor space. The methodology includes generating a correlation matrix, interpreting factor axes, plotting correlation circles, and examining individual contributions. Results provide insights into the distribution of vehicles across the factor map, supported by graphical visualizations.

Principal Component Analysis
PCA
FactoMineR
10p0
Fouille de données

This document outlines an academic workshop focused on data mining techniques used within the context of business analytics and data science. It covers principal component analysis (PCA) of vehicle data, including tasks such as correlation analysis, factor interpretation, and data visualization techniques using R packages. Practical coding examples are provided to assist in performing PCA and interpreting results.

quot
analyse
package
10p0
Business Analytics & Data Science - Atelier: Fouille de données

This document outlines an atelier focused on data mining within the Business Analytics & Data Science program for the academic year 2021-2022. It covers principal component analysis, correspondence analysis, and unsupervised classification methods. Additionally, it provides references to relevant resources such as cran.r-project.org and anaconda.com.

analyse
analytics
atelier
4p0
Atelier fouille de données

This document presents a data mining project focused on classifying road accidents in France for the year 2020. It includes data importation, variable selection, descriptive statistics, correlation analysis, PCA, and K-means clustering. The findings highlight the relationships among various factors contributing to road accidents.

quot
data
variance
11p0
Introduction à Apache Spark - Atelier Apache Spark

Apache Spark is a powerful, open-source cluster computing platform designed for processing large-scale data. It provides high-level APIs in Java, Scala, Python, and R and supports batch and stream processing while leveraging in-memory computation for enhanced performance. Spark's layered architecture includes components like Spark SQL for structured data processing, Spark Streaming for real-time analytics, Spark MLlib for machine learning, and Spark GraphX for graph computations. Despite its advantages in processing speed and scalability, Spark has limitations such as higher memory costs, l...

Apache Spark
RDD
Machine Learning MLlib
7p0
Introduction à Apache Spark - Atelier Apache Spark

This document introduces Apache Spark, a powerful open-source platform for cluster data processing. It covers the architecture, key components, and functionalities such as batch processing, stream processing, and machine learning capabilities. Apache Spark integrates well with Hadoop and provides a high-level API in multiple programming languages.

spark
apache
traitement
7p0
Introduction au Big Data - Atelier Apache Spark

The document introduces the concept of Big Data by explaining its origins, characteristics, and challenges it imposes on traditional data management systems. It discusses the core 'Vs' of Big Data (Volume, Velocity, Variety) and their consequences on system architecture, such as the need for horizontal scalability, continuous availability, and flexibility. The text also covers the CAP theorem and its implications for distributed Big Data systems, highlighting choices between consistency, availability, and partition tolerance. Solutions like distributed and scalable infrastructures are neces...

Big Data
Volume
Velocity
14p0
Introduction au Big Data - Atelier Apache Spark

This document introduces the concept of Big Data, highlighting its significance and the reasons behind its emergence. It discusses the need for new infrastructures to handle the vast amounts of data generated in today's digital age. Furthermore, it outlines the characteristics of Big Data, particularly focusing on the volume aspect and its implications.

donn
data
syst
14p0
TP 5: Compte Rendu

This document discusses the impact of adjusting the 'max_depth' parameter in a machine learning model. It highlights the problem of overfitting and underfitting based on changes to this parameter and identifies the optimal value for achieving the best mean squared error. The findings include specific error values computed for test datasets and emphasize the application of grid search to optimize 'max_depth'.

max_depth
mean_squared_error
overfiting
1p0
TP9: Utilisation des expressions régulières en R pour le traitement des données email

This document explores the application of regular expressions in R for processing email data. Initially, challenges in parsing email domains are identified, leading to the utilization of specific regular expressions to extract data systematically. Functions like 'grep', 'gsub', 'regexpr', and 'regmatches' are discussed and applied to manipulate, search, and modify email strings effectively. Results include generating a frequency table for email domains from the processed dataset.

regular expressions
R programming
data parsing
2p0
Data Processing with Regular Expressions in R

This document discusses data reading, the application of regular expressions in R, and how to extract domains from email addresses. It provides coding examples and explanations of the syntax used in regular expressions. The focus is on overcoming challenges associated with unique email formats and counting domain frequencies.

quot
email
nous
2p0
Introduction to Data Import and Manipulation in R

The document provides an overview of methods to import, read, and manipulate tabular data formats in R, such as CSV and Excel files. It discusses the use of built-in functions like `read.csv` and libraries like `readxl` for handling external datasets. Methods to access specific rows, columns, and subsets of these data frames are thoroughly explained. Additionally, the document introduces built-in datasets in R and demonstrates their usage with key functions such as `data()` and `help()`.

R programming
CSV
Excel files
3p0
TP : Génération des outputs sur SPSS

This document describes a series of exercises focused on generating and interpreting outputs using SPSS statistical software. Tasks include creating frequency tables and cross-tabulations for variables such as travel purpose, global satisfaction, and respondent demographics (e.g., gender, profession). Graphs such as histograms and pie charts are created to visualize data, and analytical insights involve calculating age statistics and average scores for key criteria (ticket price, service quality, and flight punctuality). The exercises emphasize cross-tabulation with percentages for deeper i...

SPSS
frequency tables
cross-tabulation
2p0
APRENTISSAGE AVEC PYTHON ET SICKIT LEARN

This document outlines a lab project focused on implementing the K-Nearest Neighbors (K-NN) algorithm using the MNIST dataset. It provides objectives, references, and a hands-on approach to loading and visualizing handwritten digits. Students will apply the K-NN classifier to make predictions based on the dataset.

donne
mnist
test
5p0
TP Classification

This lab focuses on developing an algorithm to classify images of different shapes (hearts, clubs, diamonds, spades) using various features such as central and Hu moments. It includes steps for segmentation, data selection, feature extraction, and testing the model using a confusion matrix. The goal is to improve the model's accuracy by selecting relevant features.

data
base64
image
8p0
Apprentissage avec Python et Scikit-Learn – TP6: Forêts aléatoires

This document provides hands-on exercises with ensemble learning methods such as bagging, random forests, and boosting using Python's Scikit-learn library. Methods like BaggingClassifier and RandomForestClassifier are explored systematically in the context of reducing variance and improving model accuracy, with practical examples applied on the Digits dataset. Readers are tasked with evaluating model performance through accuracy, variance analysis, and graphical insights. The final sections investigate the role of parameter tuning in improving weak learners and optimizing ensemble models.

ensemble learning
BaggingClassifier
RandomForestClassifier
6p0
Apprentissage avec Python et Scikit-learn: Les arbres de décision en régression

This document focuses on regression using decision trees with the Scikit-learn library. The methodology involves training a regression model on sinusoidal data with added noise and exploring model performance by adjusting parameters like max_depth. Practical exercises include modifying the noise level in data and tuning the decision tree parameters for optimized performance. Additionally, it introduces using the Diabetes dataset to evaluate mean squared error and conduct parameter tuning via grid search.

decision trees
DecisionTreeRegressor
grid search
3p0
TP Business Intelligence

The document describes a practical case study centered on designing and implementing a data warehouse for a fictitious global sports goods company, Orion. It covers the company’s organizational, product, and customer informational landscape and establishes questions that the Business Intelligence system aims to answer, focusing on performance metrics and decision support. The solution involves leveraging an operational database and preparing a star schema warehouse with dimensions and fact tables, which will be populated using ETL processes created within Talend Open Studio. Practical instr...

data warehouse
ETL processes
Talend Open Studio
13p0
Import and Manipulation of Tabular Data in R

This document provides an introductory tutorial on importing and reading tabular data formats (CSV, Excel) in R, detailing commands like `download.file`, `read.csv`, and the use of the external `readxl` library for Excel files. It explains how to access rows, columns, and subsets of data using dataframes, demonstrates built-in datasets in R, and explores additional capabilities like inspecting structure (`str`) and accessing complex calculations for data manipulation. Methods for installing and utilizing libraries are also covered comprehensively.

dataframes
read.csv
readxl
2p0

Autres ressources en intelligence artificielle et données