Intelligence artificielle et données

307 documents à télécharger gratuitement

Cours, examens, TD, TP et exercices de intelligence artificielle et données. Thèmes couverts : machine learning, apprentissage automatique, big data, data science, fouille de données.

Concours Nationaux d’Entrée aux Cycles de Formation d’Ingénieurs Session 2019

This document outlines the competition for entrance into engineering training cycles in Tunisia. It includes details about the mathematics and physics examination, along with a focus on a computer science problem related to supervised machine learning. The exam consists of problems designed to test the implementation of binary tree structures for predictive modeling.

dset
cision
repr
10p0
Pourquoi utiliser Python pour l'IA et l'apprentissage automatique?

This document highlights Python's suitability for AI and Machine Learning projects due to its simplicity, flexibility, and robust ecosystem of libraries and frameworks. Python enables developers to focus on solving ML problems rather than dealing with language complexities. It supports platform independence, efficient collaboration, and cost-effective model training. With a strong community and tools like TensorFlow, Scikit-learn, and NumPy, Python leads as the go-to programming language for AI-driven applications.

Python
Scikit-learn
TensorFlow
6p0
Apprentissage avec Python et Scikit-learn

This document introduces the implementation and performance evaluation of multi-layer perceptron classifiers using Scikit-learn. It provides guidelines for classifying the Iris flower dataset, focusing on configuring MLP architectures for optimal results. The report is inspired by resources from the Scikit-learn and Python official sites. It emphasizes deadlines and highlights external references for further support.

MLPClassifier
Scikit-learn
classification
1p0
Séparation Linéaire

This document introduces the concept of linear separation and its applications in binary classification problems. It defines the decision function g(x | w, b) and describes the conditions for classifying data points into binary categories. Multi-class classification is briefly referenced, and mathematical examples such as support vector machines (SVM) and the Iris dataset are mentioned for practical illustration. The notes emphasize the role of weights and biases in determining decision boundaries.

linear separation
g(x
w
26p0
Modélisation du Data Warehouse

The document outlines the methodology for structuring and designing data warehouses through source data analysis, employing techniques like star or snowflake schema modeling. Logical data schema are applied for efficient storage and retrieval. The work ensures optimized data organization, critical for reporting and decision-making processes. The project execution describes strategies for handling complex data relationships and transformations.

Data warehouse
Star schema
Snowflake schema
2p0
Activité 2 : Conception et Implémentation du Data Warehouse

This document outlines the steps for designing and implementing a Data Warehouse tailored to decision-makers' needs for data visualization. Key Performance Indicators (KPIs) relevant to the business, such as sales targets, customer retention rates, and production efficiency, are identified and their calculation methods are provided. The feasibility of retrieving these KPIs from existing production databases is to be assessed, and a comprehensive data dictionary for the source databases is expected. A star or constellation schema modeling approach is proposed, followed by implementation usin...

Data Warehouse
Key Performance Indicators (KPIs)
SQL Server
2p0
Conception et Implémentation du Data Warehouse

This document outlines an activity for the design and implementation of a data warehouse based on decision-maker needs for data visualization. It guides students through identifying key performance indicators (KPIs), assessing data source feasibility, creating a data dictionary, and proposing a star/constellation modeling for the data warehouse. Implementation should be done using SQL Server.

data
nombre
donne
2p0
Analyse multidimensionnelle sur les performances des athlètes

This document analyzes the performance of athletes in two sports events using multivariate data analysis. It includes 10 variables describing each athlete and focuses on deriving principal components. The eigenvalue table outlines the variance explained by each dimension, with detailed correlation and contribution values for the variables. Further analysis involves interpreting the factor map, grouping individuals, and understanding the characteristics of the identified groups.

Multivariate Analysis
Principal Component Analysis
Correlation
5p0
Petit aide-mémoire SPSS

This document serves as a quick reference guide for using SPSS. It covers data coding, transferring data from Excel to SPSS, and various SPSS functions for data analysis. Additionally, it provides step-by-step instructions on labeling variables and visualizing results.

variable
quot
variables
27p0
DEVOIR SURVEILLE: SR & F POUR BIG DATA

The document is a supervised exam focused on distributed systems and their application in big data. It evaluates the understanding of vector clocks, matrix clocks, and token-based mutual exclusion in distributed systems. The exercise requires calculating time synchronization within systems, resolving mutual exclusion scenarios, and schematizing critical section access sequences based on specific conditions. The methodologies demonstrate a combination of theoretical models and practical problem-solving in distributed environments.

distributed systems
vector clock
matrix clock
3p0
Devoir Surveillé: Systèmes Répartis et Fonctionnement pour Big Data

The document focuses on analyzing distributed systems using vector clocks and matrix clocks for synchronization. It includes exercises to calculate dependencies and mutual exclusion scenarios with token-based control. Additionally, critical section entry sequences are modeled under different configurations in a distributed setup.

distributed systems
vector clock
matrix clock
3p0
Devoir Surveillé: SR & F pour Big Data

The document contains two exercises focusing on distributed systems. The first exercise involves calculating vector and matrix clocks to track chronological events in the given distributed system. The second exercise addresses mutual exclusion using a token-based control mechanism for a distributed system with four sites, modeling scenarios for ordering critical section access based on token availability and requests.

clock synchronization
vector clock
matrix clock
3p0
Projet Décisionnel 2020/2021

This document outlines the structure and evaluation plan for the 'Projet Décisionnel' course for 2020/2021. Key focuses include selecting a business process, designing a data warehouse using tools like Power AMC, implementing ETL with SSIS, creating and querying OLAP cubes via SSAS, and conducting reporting with PowerBI. The project deliverables include interim and final solutions, evaluated through interaction, presentations, and a final defense.

Data Warehouse
Power AMC
ETL
5p0
Projet Décisionnel

This document outlines the objectives and deliverables of a data warehousing course for the academic year 2020/2021. The module covers the selection of business processes, key performance indicators, and the design and implementation of a Data Warehouse utilizing various tools. Students are required to work in pairs and deliver presentations and reports throughout the course.

number
outil
slide
5p0
M1 BADS - Devoir en Traitement des données

This document presents a statistical analysis exercise split into two parts. The first focuses on performing one-way ANOVA tests on decathlon event data to determine competition-based differences in the 100m and javelin events. The second part analyzes 16 years of financial indicators from an oil company, explaining how to construct a data set for Principal Component Analysis (PCA) in SPSS. It identifies variables and observations, evaluates inertia distribution on axes, and explores correlations and patterns in financial data. Key findings include strong negative correlations between EXP a...

ANOVA
PCA
Kaiser
3p0
Introuction au Apache Spark cours pdf

Partie 2 - Introduction Apache Spark Apache Spark - Pr sentation Apache Spark est une plateforme de traitement sur cluster g n rique. C'est un moteur de traitement libre, assurant un traitement parall le et distribu sur des donn es massives.

Artificial Intelligence
lab
spark
1p0
Introduction à Apache Spark - Atelier Apache Spark

Apache Spark is a powerful, open-source cluster computing platform designed for processing large-scale data. It provides high-level APIs in Java, Scala, Python, and R and supports batch and stream processing while leveraging in-memory computation for enhanced performance. Spark's layered architecture includes components like Spark SQL for structured data processing, Spark Streaming for real-time analytics, Spark MLlib for machine learning, and Spark GraphX for graph computations. Despite its advantages in processing speed and scalability, Spark has limitations such as higher memory costs, l...

Apache Spark
RDD
Machine Learning MLlib
7p0
Introduction à Apache Spark - Atelier Apache Spark

This document introduces Apache Spark, a powerful open-source platform for cluster data processing. It covers the architecture, key components, and functionalities such as batch processing, stream processing, and machine learning capabilities. Apache Spark integrates well with Hadoop and provides a high-level API in multiple programming languages.

spark
apache
traitement
7p0
Introduction au Big Data - Atelier Apache Spark

The document introduces the concept of Big Data by explaining its origins, characteristics, and challenges it imposes on traditional data management systems. It discusses the core 'Vs' of Big Data (Volume, Velocity, Variety) and their consequences on system architecture, such as the need for horizontal scalability, continuous availability, and flexibility. The text also covers the CAP theorem and its implications for distributed Big Data systems, highlighting choices between consistency, availability, and partition tolerance. Solutions like distributed and scalable infrastructures are neces...

Big Data
Volume
Velocity
14p0
Introduction au Big Data - Atelier Apache Spark

This document introduces the concept of Big Data, highlighting its significance and the reasons behind its emergence. It discusses the need for new infrastructures to handle the vast amounts of data generated in today's digital age. Furthermore, it outlines the characteristics of Big Data, particularly focusing on the volume aspect and its implications.

donn
data
syst
14p0
TP 5: Compte Rendu

This document discusses the impact of adjusting the 'max_depth' parameter in a machine learning model. It highlights the problem of overfitting and underfitting based on changes to this parameter and identifies the optimal value for achieving the best mean squared error. The findings include specific error values computed for test datasets and emphasize the application of grid search to optimize 'max_depth'.

max_depth
mean_squared_error
overfiting
1p0
DS Machine Learning

The document discusses decision tree learning as a predictive modeling technique commonly used in statistics, data mining, and machine learning. It explains the decision tree structure, with branches representing observations and leaves representing conclusions about target values. The strengths highlighted include intelligible knowledge representation and automatic variable selection, while weaknesses point to stability issues with small datasets and challenges in detecting variable interactions. The document emphasizes decision tree robustness and efficiency for medium-sized datasets.

decision tree learning
predictive modeling
statistics
1p0
Decision Trees: Strengths and Weaknesses

Decision tree learning is a predictive modeling approach commonly used in statistics, data mining, and machine learning. It employs a tree structure where observations determine branches and conclusions determine leaf values. Key strengths include intelligible knowledge validation, automatic variable selection, and robust performance against outliers. However, weaknesses include instability with small datasets, difficulty in identifying variable interactions, and masking of selected variable importance.

decision trees
predictive modeling
learning algorithm
1p0
Arbres de décision en apprentissage automatique

This document explores the use of decision tree learning as a predictive modeling approach within statistics and machine learning. It outlines the strengths and weaknesses of this technique, demonstrating its effectiveness and challenges in data analysis. Key points include the intelligibility of the model and its robustness against outliers, along with issues related to stability in small datasets.

quot
arbre
variables
1p0

Autres ressources en intelligence artificielle et données