Intelligence artificielle et données
307 documents à télécharger gratuitement
Cours, examens, TD, TP et exercices de intelligence artificielle et données. Thèmes couverts : machine learning, apprentissage automatique, big data, data science, fouille de données.
This document discusses the implementation of the K-Nearest Neighbors (K-NN) algorithm for classification, using the MNIST dataset, a labeled collection of handwritten digit images commonly used in machine learning for supervised learning tasks. The methodology involves data preprocessing, training/testing dataset splitting, visualization, and classifier training using Scikit-Learn's 'KNeighborsClassifier'. The K-NN model, when trained and tested, achieved a classification accuracy of approximately 98%. Additional tasks include experimenting with various values of K to compare classificatio...
This document outlines a lab project focused on implementing the K-Nearest Neighbors (K-NN) algorithm using the MNIST dataset. It provides objectives, references, and a hands-on approach to loading and visualizing handwritten digits. Students will apply the K-NN classifier to make predictions based on the dataset.
This document explores the application of linear discriminant analysis (LDA) using the famous Fisher's Iris dataset. The dataset, containing measurements of iris flowers, is analyzed with R software to classify flowers into three species—setosa, versicolor, and virginica—based on biometric features. LDA results demonstrate group means, discriminant coefficients, and visualization of species separation in the feature space. A predictive example calculates scores to classify a flower manually, complemented by similar steps implemented in Python.
This document presents the concept and applications of discriminant analysis in various fields, including medicine and finance. It describes how to separate individuals into classes based on qualitative and quantitative variables. Additionally, it discusses the methods for identifying discriminative variables and making predictions based on observations.
This document introduces the application of decision trees using the Gini Index as a measurement. A banking example is demonstrated, predicting client loan repayment success based on attribute selection. It explains step-by-step tree construction, detailizing rules extraction and key strengths such as interpretability, automatic variable selection, and robustness. Lastly, it discusses advantages like efficiency on medium datasets and drawbacks including instability on small datasets.
This document explores decision trees with an emphasis on using the Gini index for predictive modeling in a banking context. It outlines the construction and interpretation of decision trees based on client repayment success. The document also discusses the strengths and weaknesses of decision trees in data analysis.
This document provides a detailed walkthrough of calculating the Gini Index, a metric commonly used in decision tree models and statistical analysis, to assess data diversity distributions. It uses a hypothetical dataset with client demographic, account, and internet usage characteristics, segmented by variables such as age and financial indicators. The calculations focus on splitting data into subsets defined by mean account balances (low, medium, and high) and computing Gini values before and after these splits. The analysis demonstrates the reduction in impurity upon appropriate segmenta...
This document outlines a practical assignment focused on implementing Naive Bayes classification for spam filtering using Python and Scikit-learn. The assignment emphasizes the creation of a Bayesian classifier to correlate word occurrences with spam or legitimate emails (ham). Key external references include the official Scikit-learn and Python websites. The project serves as an introduction to practical applications of Naive Bayes in combating email spam, a persistent problem in digital communication.
This practical assignment introduces the use of Support Vector Machines (SVMs) for linear classification tasks using Python and the Scikit-learn library. The Iris dataset serves as the example, focusing on its first two attributes, with half used for training and half for testing. The methodology includes training a Linear SVM, tuning the regularization parameter 'C', and optimizing performance by examining all four attributes of the dataset. The assignment includes tasks to analyze model accuracy, decision boundaries, and potential improvements for the method.
This document provides a comprehensive overview of Support Vector Machines (SVMs) and kernel methods. It covers linear and non-linear SVMs, including the mathematical formulation and optimization techniques. Topics include soft margins for handling non-linearly separable data, kernel tricks for higher dimensional space mapping, and dual optimization approaches. Practical applications and software tools such as Scikit-learn, LibSVM, and Torch are also discussed.
This document covers the principles of Support Vector Machines (SVM), including concepts of linear and non-linear separability. It discusses the mathematical foundations and applications of SVM in various contexts. Additionally, it explores the optimization techniques necessary for effective implementation of SVM.
This document outlines a hands-on activity to perform simple linear regression analysis using R and Scikit-learn. Key tasks include data import, descriptive statistics computation, correlation analysis, and building a regression model to explain 'price' as a function of 'horsepower'. Additional steps involve visualization with scatter plots, removal of outliers, calculation of regression coefficients and R² metrics, and making predictions based on the regression model. The activity integrates practical use of scientific tools for comprehensive data analysis.
Hibernate est un framework Java qui simplifie le développement d'applications Java pour interagir avec la base de données. Il s'agit d'un outil open source, léger, ORM (Object Relational Mapping).
Il est largement reconnu que l'intelligence artificielle jouera un r le croissant dans nos vies. mesure que la technologie volue, le besoin de dipl m s poss dant des connaissances sp cialis es augmente.
This lab focuses on developing an algorithm to classify images of different shapes (hearts, clubs, diamonds, spades) using various features such as central and Hu moments. It includes steps for segmentation, data selection, feature extraction, and testing the model using a confusion matrix. The goal is to improve the model's accuracy by selecting relevant features.
This document provides hands-on exercises with ensemble learning methods such as bagging, random forests, and boosting using Python's Scikit-learn library. Methods like BaggingClassifier and RandomForestClassifier are explored systematically in the context of reducing variance and improving model accuracy, with practical examples applied on the Digits dataset. Readers are tasked with evaluating model performance through accuracy, variance analysis, and graphical insights. The final sections investigate the role of parameter tuning in improving weak learners and optimizing ensemble models.
This document outlines advanced ensemble learning methods such as Bagging, Random Forests, and Boosting, focusing on reduction of variance and improvement in model performance. High-variance estimators are discussed, with Bagging reducing variance through bootstrap sampling and averaging, and Random Forests further decorrelating predictors by introducing randomness in attribute selection during tree construction. Boosting enhances weak learners by iteratively minimizing weighted error based on previous performance, effectively generating a robust final model. Empirical results, parameter tu...
In this document, data mining is introduced as an interdisciplinary field combining statistics, artificial intelligence, and database technologies to extract valuable information from large datasets. It explores methodologies such as KDD, SEMMA, and CRISP-DM, discussing their application to decision-making and fraud detection. Domains of application range from financial scoring to medical diagnostics and predictive analysis. The material also delves into the preparation of data and two primary categories of techniques: supervised learning for predictive tasks and unsupervised learning for p...
The document introduces supervised classification methods from the perspective of automatic learning techniques. It explains the inductive process of deriving classification rules from labeled examples, contrasting symbolic (e.g., decision trees) and adaptive (e.g., neural networks, genetic algorithms) methods. An emphasis is placed on decision tree algorithms, their interpretability, and their application in various domains, such as medical diagnostics or customer behavior prediction. Techniques like entropy and Gini measures for assessing classification purity are discussed to highlight d...
Ce document examine les méthodes de classification supervisée en utilisant des systèmes d'apprentissage. Il compare les approches symboliques, comme les arbres de décision, aux méthodes non paramétriques et adaptatives, telles que les réseaux de neurones. L'accent est mis sur l'importance d'une procédure de classification interprétable pour des applications comme le diagnostic médical.
This document focuses on regression using decision trees with the Scikit-learn library. The methodology involves training a regression model on sinusoidal data with added noise and exploring model performance by adjusting parameters like max_depth. Practical exercises include modifying the noise level in data and tuning the decision tree parameters for optimized performance. Additionally, it introduces using the Diabetes dataset to evaluate mean squared error and conduct parameter tuning via grid search.
This document explores decision trees in machine learning, discussing their hierarchical representation, construction methodology, and practical applications in classification and regression tasks. It delves into key algorithms such as ID3, C4.5, and CART, explaining their entropy-based split criteria and information gain processes. Implementation challenges like overfitting and missing data handling are addressed through pruning and surrogate splits. The document concludes with advanced topics including ensemble methods like bagging and random forests for improved generalization.
The document focuses on decision trees, a machine learning technique for data exploration and decision-making. It explains various concepts including the motivation, definitions, and methodologies for classification and regression tasks. It compares key algorithms such as ID3, C4.5, C5.0, and CART, discussing tree-building principles, entropy-based splitting, and validation methods. Extensions like bagging and random forests are also covered, emphasizing their advantages for improving decision-making accuracy.
Ce cours aborde les arbres de décision, leurs définitions, motivations et applications. Il couvre également l'apprentissage automatique à travers la classification et la régression, ainsi que des implémentations pratiques. Des techniques avancées, telles que le bagging et les forêts aléatoires, sont aussi discutées.















