Travaux pratiques (TP) - Intelligence artificielle et données
22 documents à télécharger gratuitement
Travaux pratiques (TP) de intelligence artificielle et données, partagés par des étudiants et des enseignants. Thèmes couverts : machine learning, apprentissage automatique, big data, data science, fouille de données.
This document explores unsupervised classification methods, focusing primarily on two types: hierarchical and non-hierarchical methods. It details the K-means algorithm, explaining its principles, convergence criteria, and the role of intra-class and inter-class variance. Hierarchical clustering is also discussed, describing dendrograms and aggregation criteria such as single linkage, complete linkage, and variance-based methods. Additionally, the document provides an overview of mixed classification, combining K-means partitioning with hierarchical clustering for a more refined analysis.
This document outlines methods of unsupervised classification, focusing on automatic classification techniques. It discusses the concepts of dissimilarity and inertia as they relate to grouping individuals based on their characteristics. Additionally, it presents the K-means algorithm for classifying data.
This document outlines a course workshop on data mining as part of the Business Analytics & Data Science master's program. It includes a section on principal component analysis (PCA) applied to vehicle data with specific tasks for analysis. The focus is on interpreting correlations and distributions of vehicles using PCA techniques.
The document outlines a practical workshop for Principal Component Analysis (PCA) applied to six continuous variables recorded from data on 20 cars. Using R programming and the FactoMineR and factoextra packages, correlations between variables are analyzed to identify patterns and influences in the factor space. The methodology includes generating a correlation matrix, interpreting factor axes, plotting correlation circles, and examining individual contributions. Results provide insights into the distribution of vehicles across the factor map, supported by graphical visualizations.
This document outlines an academic workshop focused on data mining techniques used within the context of business analytics and data science. It covers principal component analysis (PCA) of vehicle data, including tasks such as correlation analysis, factor interpretation, and data visualization techniques using R packages. Practical coding examples are provided to assist in performing PCA and interpreting results.
This document outlines an atelier focused on data mining within the Business Analytics & Data Science program for the academic year 2021-2022. It covers principal component analysis, correspondence analysis, and unsupervised classification methods. Additionally, it provides references to relevant resources such as cran.r-project.org and anaconda.com.
This document presents a data mining project focused on classifying road accidents in France for the year 2020. It includes data importation, variable selection, descriptive statistics, correlation analysis, PCA, and K-means clustering. The findings highlight the relationships among various factors contributing to road accidents.
Apache Spark is a powerful, open-source cluster computing platform designed for processing large-scale data. It provides high-level APIs in Java, Scala, Python, and R and supports batch and stream processing while leveraging in-memory computation for enhanced performance. Spark's layered architecture includes components like Spark SQL for structured data processing, Spark Streaming for real-time analytics, Spark MLlib for machine learning, and Spark GraphX for graph computations. Despite its advantages in processing speed and scalability, Spark has limitations such as higher memory costs, l...
This document introduces Apache Spark, a powerful open-source platform for cluster data processing. It covers the architecture, key components, and functionalities such as batch processing, stream processing, and machine learning capabilities. Apache Spark integrates well with Hadoop and provides a high-level API in multiple programming languages.
The document introduces the concept of Big Data by explaining its origins, characteristics, and challenges it imposes on traditional data management systems. It discusses the core 'Vs' of Big Data (Volume, Velocity, Variety) and their consequences on system architecture, such as the need for horizontal scalability, continuous availability, and flexibility. The text also covers the CAP theorem and its implications for distributed Big Data systems, highlighting choices between consistency, availability, and partition tolerance. Solutions like distributed and scalable infrastructures are neces...
This document introduces the concept of Big Data, highlighting its significance and the reasons behind its emergence. It discusses the need for new infrastructures to handle the vast amounts of data generated in today's digital age. Furthermore, it outlines the characteristics of Big Data, particularly focusing on the volume aspect and its implications.
This document discusses the impact of adjusting the 'max_depth' parameter in a machine learning model. It highlights the problem of overfitting and underfitting based on changes to this parameter and identifies the optimal value for achieving the best mean squared error. The findings include specific error values computed for test datasets and emphasize the application of grid search to optimize 'max_depth'.
This document explores the application of regular expressions in R for processing email data. Initially, challenges in parsing email domains are identified, leading to the utilization of specific regular expressions to extract data systematically. Functions like 'grep', 'gsub', 'regexpr', and 'regmatches' are discussed and applied to manipulate, search, and modify email strings effectively. Results include generating a frequency table for email domains from the processed dataset.
This document discusses data reading, the application of regular expressions in R, and how to extract domains from email addresses. It provides coding examples and explanations of the syntax used in regular expressions. The focus is on overcoming challenges associated with unique email formats and counting domain frequencies.
The document provides an overview of methods to import, read, and manipulate tabular data formats in R, such as CSV and Excel files. It discusses the use of built-in functions like `read.csv` and libraries like `readxl` for handling external datasets. Methods to access specific rows, columns, and subsets of these data frames are thoroughly explained. Additionally, the document introduces built-in datasets in R and demonstrates their usage with key functions such as `data()` and `help()`.
This document describes a series of exercises focused on generating and interpreting outputs using SPSS statistical software. Tasks include creating frequency tables and cross-tabulations for variables such as travel purpose, global satisfaction, and respondent demographics (e.g., gender, profession). Graphs such as histograms and pie charts are created to visualize data, and analytical insights involve calculating age statistics and average scores for key criteria (ticket price, service quality, and flight punctuality). The exercises emphasize cross-tabulation with percentages for deeper i...
This document outlines a lab project focused on implementing the K-Nearest Neighbors (K-NN) algorithm using the MNIST dataset. It provides objectives, references, and a hands-on approach to loading and visualizing handwritten digits. Students will apply the K-NN classifier to make predictions based on the dataset.
This lab focuses on developing an algorithm to classify images of different shapes (hearts, clubs, diamonds, spades) using various features such as central and Hu moments. It includes steps for segmentation, data selection, feature extraction, and testing the model using a confusion matrix. The goal is to improve the model's accuracy by selecting relevant features.
This document provides hands-on exercises with ensemble learning methods such as bagging, random forests, and boosting using Python's Scikit-learn library. Methods like BaggingClassifier and RandomForestClassifier are explored systematically in the context of reducing variance and improving model accuracy, with practical examples applied on the Digits dataset. Readers are tasked with evaluating model performance through accuracy, variance analysis, and graphical insights. The final sections investigate the role of parameter tuning in improving weak learners and optimizing ensemble models.
This document focuses on regression using decision trees with the Scikit-learn library. The methodology involves training a regression model on sinusoidal data with added noise and exploring model performance by adjusting parameters like max_depth. Practical exercises include modifying the noise level in data and tuning the decision tree parameters for optimized performance. Additionally, it introduces using the Diabetes dataset to evaluate mean squared error and conduct parameter tuning via grid search.
The document describes a practical case study centered on designing and implementing a data warehouse for a fictitious global sports goods company, Orion. It covers the company’s organizational, product, and customer informational landscape and establishes questions that the Business Intelligence system aims to answer, focusing on performance metrics and decision support. The solution involves leveraging an operational database and preparing a star schema warehouse with dimensions and fact tables, which will be populated using ETL processes created within Talend Open Studio. Practical instr...
This document provides an introductory tutorial on importing and reading tabular data formats (CSV, Excel) in R, detailing commands like `download.file`, `read.csv`, and the use of the external `readxl` library for Excel files. It explains how to access rows, columns, and subsets of data using dataframes, demonstrates built-in datasets in R, and explores additional capabilities like inspecting structure (`str`) and accessing complex calculations for data manipulation. Methods for installing and utilizing libraries are also covered comprehensively.


















