Intelligence artificielle et données
307 documents à télécharger gratuitement
Cours, examens, TD, TP et exercices de intelligence artificielle et données. Thèmes couverts : machine learning, apprentissage automatique, big data, data science, fouille de données.
The document provides an in-depth exploration of classification techniques within data mining, focusing on methods such as k-means and hierarchical clustering. It describes the definition, key properties, and applications of classification, emphasizing the process of grouping objects into homogenous clusters. Various evaluation criteria and methods for measuring classification quality, such as interclass and intraclass inertia, are explained. Practical examples, including a step-by-step implementation of k-means clustering, are provided to illustrate the methodology.
This document involves the application of a Principal Component Analysis (PCA) on the performance data of 31 students from the STID1 program, covering four coursework grades: Informatics, Algorithms, Mathematical Foundations, and Mathematical Techniques. The analysis was performed using the SAS software, focusing on interpreting statistical descriptors like means and standard deviations, building a correlation matrix, and diagonalizing it to extract eigenvalues and eigenvectors. The core findings highlight the number of axes needed for meaningful visualization, significant contributors to t...
The document summarizes the principles and applications of correspondence analysis methods (AFC and ACM). It highlights their advantages, such as transforming qualitative variables into quantitative ones, handling nonlinear relationships, and visualizing variable dependencies and patterns. Case studies like store client segmentation are leveraged to compute metrics such as eigenvalues and contributions, with detailed analysis for dimensions and axes. The explanation includes the retention of significant axes based on inertia distribution, with practical insights into variable contributions...
The document provides an in-depth tutorial on Local Outlier Factor (LOF), a method to identify anomalies in datasets by comparing the local density of a point to its nearest neighbors. It highlights key concepts like outlier detection and novelty detection, mathematical frameworks (e.g., Mahalanobis distance and reachability density), and its practical implementation in R. The tutorial addresses challenges, such as choosing the optimal parameter k, and underscores the method’s advantages and limitations in detecting anomalies robustly, especially in clustered data.
This document explores fundamental factorial analysis techniques including PCA, FCA, and MCA, as preliminary methods for multivariate analysis. PCA focuses on quantitative variables and projects data into lower-dimensional subspaces while preserving distances between individuals. FCA is applied to contingency tables to analyze relationships between qualitative variables, and MCA generalizes FCA to account for more than two qualitative variables using disjunctive tables and Burt tables. The document outlines mathematical foundations, steps for performing factorial analyses, and graphical int...
The document provides solutions to a series of exercises focused on data mining concepts such as centroid calculation, Euclidean distance, and inertia determination relative to point clouds. Methodologies such as Min-Max normalization, Z-score normalization, and decimal scaling are detailed, with numerical examples provided. The core mathematical computations and formulas used in data processing are demonstrated, emphasizing practical applications in data normalization and cluster analysis. The findings include specific worked-out values and examples indicative of their real-world usage in...
This document thoroughly explores the methodologies and multiple phases of data mining (DM), including defining study objectives, data collection, preprocessing, and predictive modeling. Emphasis is placed on data preparation methods such as handling missing or extreme values, detecting multicollinearity, and evaluating data distribution characteristics through techniques like normalization and discretization. Statistical tools for detecting anomalies and variable relationships, along with advanced analysis techniques like regression and variance analysis, are discussed. Additionally, sampl...
This document reviews the functionality, creation, and advanced utilization of PivotTables (Tableaux Croisés Dynamiques, TCD) in Excel. Part 1 introduces foundational concepts including TCD creation, formatting, and statistical computation. Part 2 explores data source management such as complex multi-sheet integration and relational modeling using advanced Excel features like Queries and Models. Emphasis is placed on organizing data sources effectively and leveraging advanced formula functions within TCD frameworks for dynamic analysis. It draws upon the author's extensive professional inst...
This document is an academic exam covering Natural Language Processing (NLP) methods. It assesses Web scraping techniques using various Python libraries, and preprocesses text data by addressing issues like special characters and punctuations using NLTK and regex. Additionally, it explains representing data through the TF-IDF method to identify keyword relevance. Finally, it explores textual classification techniques, likely for categorization of textual datasets.
The document presents an examination structure focused on Natural Language Processing for a Master's level course. It includes tasks such as web scraping to detect false announcements, data preprocessing following the CRIS methodology, word embedding using TF-IDF, and text classification with the Naive Bayes algorithm. Methodology for web scraping, preprocessing collected datasets, calculating TF-IDF, and building a machine learning model for detecting false announcements is elaborated through practical sections.
This document provides a detailed exam for a Master's level course on Natural Language Processing, focusing on web scraping, data preprocessing, word embedding using TF-IDF, and text classification using the Naive Bayes algorithm. The exam consists of multiple parts addressing the collection of data, preprocessing techniques, representation of data, and classification methodologies.
This document outlines a comprehensive course on natural language processing (NLP) using artificial intelligence tools. It covers advanced topics such as corpus construction and text exploration, statistical and neural approaches to word representation, sentiment analysis, topic modeling, and chatbot creation. The course emphasizes theoretical understanding, practical exercises, and application development, leveraging modern tools like Python, TensorFlow, and NLP libraries. Practical workshops focus on real-world applications, such as text classification, sentiment modeling, and web-based c...
This document introduces the core concepts of Recurrent Neural Networks (RNNs) and their limitations, such as gradient vanishing and explosion. It discusses advanced architectures like GRUs and LSTMs that incorporate gating mechanisms to address these challenges. Methodological insights include the use of Backpropagation Through Time and initialization techniques to stabilize training. The text also highlights the decline in LSTM usage in favor of transformers in modern NLP applications.
This document details fundamental optimization techniques in deep learning, beginning with gradient descent and its iterative parameter update mechanism. It further explains stochastic gradient descent, highlighting its efficiency and advantages in avoiding shallow local minima due to its stochastic nature. The concept of mini-batching is introduced to balance computational efficiency and noise reduction. Lastly, the momentum method is explored, showing how momentum aids in faster convergence by leveraging past gradients. Practical tuning advice for hyperparameters like learning rate and mo...
Le cours traite des techniques d'optimisation appliquées au deep learning, avec un accent sur la méthode de descente de gradient et ses variantes. Il aborde les défis liés à la sélection de la taille du pas optimales pour garantir la convergence sans divergence. De plus, il explore la descente de gradient stochastique et ses avantages par rapport à la descente de gradient classique.
This document provides a foundational overview of linear algebra concepts applied in neural networks, focusing on transformations and vector alignments. It extends these fundamentals to convolutions, particularly in audio data analysis, using properties like stationarity and locality to optimize calculations. Key methodologies include affine transformations, alignment principles, and weight-sharing techniques in convolution layers. The document concludes with insights into the sparsity of large matrices and their kernel representation.
This course covers the basics of linear algebra in the context of neural networks, focusing on transformations and alignments of input vectors. It extends linear algebra to convolutions, particularly in the analysis of audio data, while introducing key properties like stationarity and locality. Students will learn how to simplify complex data representations through kernel sharing and the effects of dimensionality.
This document focuses on the automatic classification of 29 varieties of cheese based on their nutritional properties using two clustering techniques: Hierarchical Ascendant Classification (CAH) and k-Means. It provides a detailed methodology that includes loading and analyzing the dataset, detecting optimal class numbers, and interpreting clustering results using univariate, multivariate, and principal component analysis. The results demonstrate that cheese groups are primarily defined by lipid and protein content, and further analysis excluding an outlier group—the 'fresh cheeses'—yields...
Ce document décrit une démarche de classification automatique d'un ensemble de fromages basé sur leurs propriétés nutritives. Deux approches seront utilisées : la classification ascendante hiérarchique et la méthode des k-Means. Ce guide offre un aperçu pratique pour utiliser Python dans le contexte de la classification automatique.
This document explores unsupervised classification methods, focusing primarily on two types: hierarchical and non-hierarchical methods. It details the K-means algorithm, explaining its principles, convergence criteria, and the role of intra-class and inter-class variance. Hierarchical clustering is also discussed, describing dendrograms and aggregation criteria such as single linkage, complete linkage, and variance-based methods. Additionally, the document provides an overview of mixed classification, combining K-means partitioning with hierarchical clustering for a more refined analysis.
This document outlines methods of unsupervised classification, focusing on automatic classification techniques. It discusses the concepts of dissimilarity and inertia as they relate to grouping individuals based on their characteristics. Additionally, it presents the K-means algorithm for classifying data.
Ce document traite de l'analyse en composantes principales appliquée à un ensemble de données sur des véhicules. Il couvre les étapes d'importation des données, des statistiques descriptives, de la standardisation, de la création de la matrice de corrélation, et enfin de l'application de l'ACP. Les résultats montrent diverses corrélations entre les caractéristiques des voitures.
This document provides a practical demonstration of Correspondence Analysis (AFC) using R. A contingency table is created from qualitative variables, followed by row and column profile analysis using proportions. Chi-square tests are conducted to assess the independence of variables, and theoretical and empirical contingency tables are computed. The correspondence analysis (CA) is performed with two principal components for data visualization and interpretation.
This document presents a practical example of Correspondence Analysis (AFC) using R. It includes various steps to analyze a contingency table between two qualitative variables. The methods demonstrated include chi-squared tests and the calculation of profile tables.






















