Intelligence artificielle et données
307 documents à télécharger gratuitement
Cours, examens, TD, TP et exercices de intelligence artificielle et données. Thèmes couverts : machine learning, apprentissage automatique, big data, data science, fouille de données.
This document outlines the competition for entrance into engineering training cycles in Tunisia. It includes details about the mathematics and physics examination, along with a focus on a computer science problem related to supervised machine learning. The exam consists of problems designed to test the implementation of binary tree structures for predictive modeling.
This document highlights Python's suitability for AI and Machine Learning projects due to its simplicity, flexibility, and robust ecosystem of libraries and frameworks. Python enables developers to focus on solving ML problems rather than dealing with language complexities. It supports platform independence, efficient collaboration, and cost-effective model training. With a strong community and tools like TensorFlow, Scikit-learn, and NumPy, Python leads as the go-to programming language for AI-driven applications.
This document introduces the implementation and performance evaluation of multi-layer perceptron classifiers using Scikit-learn. It provides guidelines for classifying the Iris flower dataset, focusing on configuring MLP architectures for optimal results. The report is inspired by resources from the Scikit-learn and Python official sites. It emphasizes deadlines and highlights external references for further support.
This document introduces the concept of linear separation and its applications in binary classification problems. It defines the decision function g(x | w, b) and describes the conditions for classifying data points into binary categories. Multi-class classification is briefly referenced, and mathematical examples such as support vector machines (SVM) and the Iris dataset are mentioned for practical illustration. The notes emphasize the role of weights and biases in determining decision boundaries.
The document outlines the methodology for structuring and designing data warehouses through source data analysis, employing techniques like star or snowflake schema modeling. Logical data schema are applied for efficient storage and retrieval. The work ensures optimized data organization, critical for reporting and decision-making processes. The project execution describes strategies for handling complex data relationships and transformations.
This document outlines the steps for designing and implementing a Data Warehouse tailored to decision-makers' needs for data visualization. Key Performance Indicators (KPIs) relevant to the business, such as sales targets, customer retention rates, and production efficiency, are identified and their calculation methods are provided. The feasibility of retrieving these KPIs from existing production databases is to be assessed, and a comprehensive data dictionary for the source databases is expected. A star or constellation schema modeling approach is proposed, followed by implementation usin...
This document outlines an activity for the design and implementation of a data warehouse based on decision-maker needs for data visualization. It guides students through identifying key performance indicators (KPIs), assessing data source feasibility, creating a data dictionary, and proposing a star/constellation modeling for the data warehouse. Implementation should be done using SQL Server.
This document analyzes the performance of athletes in two sports events using multivariate data analysis. It includes 10 variables describing each athlete and focuses on deriving principal components. The eigenvalue table outlines the variance explained by each dimension, with detailed correlation and contribution values for the variables. Further analysis involves interpreting the factor map, grouping individuals, and understanding the characteristics of the identified groups.
This document serves as a quick reference guide for using SPSS. It covers data coding, transferring data from Excel to SPSS, and various SPSS functions for data analysis. Additionally, it provides step-by-step instructions on labeling variables and visualizing results.
The document is a supervised exam focused on distributed systems and their application in big data. It evaluates the understanding of vector clocks, matrix clocks, and token-based mutual exclusion in distributed systems. The exercise requires calculating time synchronization within systems, resolving mutual exclusion scenarios, and schematizing critical section access sequences based on specific conditions. The methodologies demonstrate a combination of theoretical models and practical problem-solving in distributed environments.
The document focuses on analyzing distributed systems using vector clocks and matrix clocks for synchronization. It includes exercises to calculate dependencies and mutual exclusion scenarios with token-based control. Additionally, critical section entry sequences are modeled under different configurations in a distributed setup.
The document contains two exercises focusing on distributed systems. The first exercise involves calculating vector and matrix clocks to track chronological events in the given distributed system. The second exercise addresses mutual exclusion using a token-based control mechanism for a distributed system with four sites, modeling scenarios for ordering critical section access based on token availability and requests.
This document outlines the structure and evaluation plan for the 'Projet Décisionnel' course for 2020/2021. Key focuses include selecting a business process, designing a data warehouse using tools like Power AMC, implementing ETL with SSIS, creating and querying OLAP cubes via SSAS, and conducting reporting with PowerBI. The project deliverables include interim and final solutions, evaluated through interaction, presentations, and a final defense.
This document outlines the objectives and deliverables of a data warehousing course for the academic year 2020/2021. The module covers the selection of business processes, key performance indicators, and the design and implementation of a Data Warehouse utilizing various tools. Students are required to work in pairs and deliver presentations and reports throughout the course.
This document presents a statistical analysis exercise split into two parts. The first focuses on performing one-way ANOVA tests on decathlon event data to determine competition-based differences in the 100m and javelin events. The second part analyzes 16 years of financial indicators from an oil company, explaining how to construct a data set for Principal Component Analysis (PCA) in SPSS. It identifies variables and observations, evaluates inertia distribution on axes, and explores correlations and patterns in financial data. Key findings include strong negative correlations between EXP a...
Partie 2 - Introduction Apache Spark Apache Spark - Pr sentation Apache Spark est une plateforme de traitement sur cluster g n rique. C'est un moteur de traitement libre, assurant un traitement parall le et distribu sur des donn es massives.
Apache Spark is a powerful, open-source cluster computing platform designed for processing large-scale data. It provides high-level APIs in Java, Scala, Python, and R and supports batch and stream processing while leveraging in-memory computation for enhanced performance. Spark's layered architecture includes components like Spark SQL for structured data processing, Spark Streaming for real-time analytics, Spark MLlib for machine learning, and Spark GraphX for graph computations. Despite its advantages in processing speed and scalability, Spark has limitations such as higher memory costs, l...
This document introduces Apache Spark, a powerful open-source platform for cluster data processing. It covers the architecture, key components, and functionalities such as batch processing, stream processing, and machine learning capabilities. Apache Spark integrates well with Hadoop and provides a high-level API in multiple programming languages.
The document introduces the concept of Big Data by explaining its origins, characteristics, and challenges it imposes on traditional data management systems. It discusses the core 'Vs' of Big Data (Volume, Velocity, Variety) and their consequences on system architecture, such as the need for horizontal scalability, continuous availability, and flexibility. The text also covers the CAP theorem and its implications for distributed Big Data systems, highlighting choices between consistency, availability, and partition tolerance. Solutions like distributed and scalable infrastructures are neces...
This document introduces the concept of Big Data, highlighting its significance and the reasons behind its emergence. It discusses the need for new infrastructures to handle the vast amounts of data generated in today's digital age. Furthermore, it outlines the characteristics of Big Data, particularly focusing on the volume aspect and its implications.
This document discusses the impact of adjusting the 'max_depth' parameter in a machine learning model. It highlights the problem of overfitting and underfitting based on changes to this parameter and identifies the optimal value for achieving the best mean squared error. The findings include specific error values computed for test datasets and emphasize the application of grid search to optimize 'max_depth'.
The document discusses decision tree learning as a predictive modeling technique commonly used in statistics, data mining, and machine learning. It explains the decision tree structure, with branches representing observations and leaves representing conclusions about target values. The strengths highlighted include intelligible knowledge representation and automatic variable selection, while weaknesses point to stability issues with small datasets and challenges in detecting variable interactions. The document emphasizes decision tree robustness and efficiency for medium-sized datasets.
Decision tree learning is a predictive modeling approach commonly used in statistics, data mining, and machine learning. It employs a tree structure where observations determine branches and conclusions determine leaf values. Key strengths include intelligible knowledge validation, automatic variable selection, and robust performance against outliers. However, weaknesses include instability with small datasets, difficulty in identifying variable interactions, and masking of selected variable importance.
This document explores the use of decision tree learning as a predictive modeling approach within statistics and machine learning. It outlines the strengths and weaknesses of this technique, demonstrating its effectiveness and challenges in data analysis. Key points include the intelligibility of the model and its robustness against outliers, along with issues related to stability in small datasets.














