Intelligence artificielle et données
307 documents à télécharger gratuitement
Cours, examens, TD, TP et exercices de intelligence artificielle et données. Thèmes couverts : machine learning, apprentissage automatique, big data, data science, fouille de données.
This document introduces clustering methods, focusing on hierarchical classification and partitioning. It discusses the importance of similarity metrics in clustering, particularly the choice between Euclidean and Manhattan distances. The document also outlines the steps involved in hierarchical clustering and the concept of nested partitions.
This document presents an example of hierarchical classification on principal components (HCPC) using climate data from various European capitals. It aims to categorize capitals based on their temperature similarities throughout the year, utilizing statistical metrics and graphical representations. The analysis employs principal component analysis and clustering to unveil insights about data structure.
This document discusses various clustering methodologies, including hierarchical clustering and partitioning methods. It also emphasizes the importance of similarity measures in clustering, presenting metrics such as Euclidean and Manhattan distances. The text provides a detailed overview of hierarchical agglomerative clustering, highlighting its process and applications.
This thesis focuses on developing a decision-support system within a banking context, specifically for the Banque Internationale Arabe de Tunisie (BIAT). It employs modern Business Intelligence practices to design dashboards for the aggregation, normalization, and visualization of data. Using Ralph Kimball's methodology, the data warehouse architecture was meticulously structured to meet functional and non-functional requirements. The implementation achieved real-time data insights critical for the bank's marketing and digital development divisions, enabling enhanced decision-making and ope...
This document outlines the analysis, design, and implementation phases of building a business intelligence solution for Société Tunisienne d’Assurances et de Réassurances (STAR). It employs Scrum methodology tailored to business intelligence projects, creating data marts for monitoring production and claims. Data was extracted, transformed, and loaded (ETL) into a warehouse, followed by creating insightful dashboards and reports using Power BI. The solution improves operational reporting, performance tracking, and enables faster decision-making processes.
This document highlights the development of a decision support system for the Arab Tunisian Bank (ATB). Agile methodologies, particularly SCRUM, were employed in conjunction with Business Intelligence tools to create an optimized DataWarehouse and Datamart for analyzing client risk and asset classifications. The solution involved designing conceptual models, exploring ROLAP, MOLAP, and HOLAP approaches, and implementing Extract-Transform-Load (ETL) processes. The outcomes included enhanced data management and reporting capabilities for better financial risk assessment and client monitoring.
This report documents the development of a project aimed at generating calculations of mobile data consumption using advanced tools and methodologies. The authors evaluate existing organizational practices in OOREDOO Tunisie, propose solutions based on SCRUM BI methodology, and employ a robust Data Warehouse using tools such as Talend and Oracle Database. Key outcomes include cleaning, transforming, and integrating large data into an optimized structure, culminating in the creation of interactive dashboards for visualizing subscriber data and mobile data consumption trends.
This document details the development of a solution for estimating the prices of used cars in Tunisia. It follows an agile methodology using SCRUM, incorporating machine learning for price prediction. Core tasks include web scraping from various Tunisian car advertisement websites, data cleaning, and the implementation of linear regression and gradient regression models. The project concludes with the centralization of data in a SQL database and the development of a web interface for user interaction.
This document presents a final year project focused on the study, design, and implementation of a business intelligence dashboard system using Big Data technologies. The project uses tools like Apache Hadoop, Spark, Hive, and Arcadia Data to conceptualize and create dashboards from empirical data collected via probes (Astellia and Otarie). Key methodologies like Agile, GIMSI, and Balanced Scorecard were adopted. The result demonstrates a scalable and functional data visualization solution, offering performance indicators and insights tailored to the needs of the host organization, Orange Tu...
This document details the implementation of a Business Intelligence (BI) tool within the Odoo ERP system. Using the SCRUM methodology, it involves the design and development of various data warehouses to address distinct use cases: sales, purchasing, and stock management. The report also covers the integration of PostgreSQL for data storage and manipulation, and Power BI for visualization. Key findings include the establishment of dynamic dashboards enabling efficient data-driven decision-making for the company WAKAHAW.
These lecture notes introduce the concept of linear separation and its applications in classification problems, with mathematical formalism using the decision function g(x|w,b). The input data is described as pairs of feature vectors and class labels, with separation determined by the sign of g(x|w,b). The document also briefly mentions generalizing the method to multi-class problems and includes an example application using Support Vector Machines (SVM). Key constraints for optimization are provided to ensure functional margins.
The document is an exam for a Master's program in Business Analytics and Data Science at the Université Virtuelle de Tunis, primarily covering Machine Learning. It consists of problems focusing on the differentiation of data mining methods, discriminant analysis, and predictive modeling for employee retention. Additionally, it examines regression analyses on business investment, interpreting parameters and improving models using explanatory variables. The exam applies analytical techniques and mathematical formulations such as ordinary least squares and decision trees.
This document is the exam for the Machine Learning course as part of the Master's program in Business Analytics and Data Science at the Université Virtuelle de Tunis. It comprises various exercises evaluating theoretical knowledge and practical applications related to data mining and classification methods. Key topics include the differences between descriptive and predictive methods, classification techniques from artificial intelligence, and regression analysis.
This document outlines the third tutorial (TD no. 3) for a Data Mining course. It focuses on the k-Nearest Neighbors (kPPV) method, applying it to examples such as 'Playing Tennis' and text classification. The tutorial includes measures of dissimilarity and classification techniques using kPPV.
The document delves into the use of Correspondence Analysis (AFC) as a method within multidimensional data analysis. It emphasizes reducing dimensionality while preserving data correspondence. The methodology involves calculating inertia using eigenvalues derived from statistical observations, along with visualizing results in reduced axes. An ANOVA test and Fisher statistic are applied to determine the relevance of variables, further refining models aimed at predicting solvency or preferences.
Ce document d crit un programme d valuation de quatre mod les de machine learning. Il pr sente les tapes de pr paration, formation et valuation des mod les, ainsi qu'une analyse de la temporalit de chaque tape. Les r sultats incluent l'ex cution s quentielle des t ches et les optimisations possibles de l'ordonnancement.
The report introduces a Spark-based solution for processing retail sales data through MapReduce principles by leveraging Resilient Distributed Datasets (RDDs). A virtual environment was set up using VMware with Ubuntu 20.04, Docker, Hadoop, and Spark. The project involves data extraction, transformation using Spark RDDs, and execution of a MapReduce algorithm to count word occurrences in sales data. Results are verified through Spark's interface, showcasing a hands-on application of data processing with Spark.
This project demonstrates the implementation of a MapReduce methodology using Apache Spark by processing sales data from a text file. The environment setup involved configuring a virtual machine with Ubuntu 20.04, installing required software like Apache Hadoop, Spark, and Docker, and ensuring system compatibility. A Spark-based code was developed to analyze sales data, leveraging RDD transformations including map and reduce functions. The project successfully showcased foundational knowledge of Spark, Big Data tools, and data processing techniques.
This document is an examination covering foundational and advanced concepts in Big Data Analytics and Deep Learning. The Big Data section focuses on tools and frameworks such as Hadoop, Hive, Sqoop, Flume, and Spark, emphasizing technical commands, configurations, and comparisons. The Deep Learning section assesses conceptual understanding of neural networks, activation functions, architectures like CNN and RNN, training methodologies like gradient descent, and applications such as autoencoders and hyperparameter tuning. Methodologies of data processing, model optimization, and their advant...
The exam tests knowledge across key concepts in big data analytics, including an understanding of Hadoop 2 architecture, Elasticsearch clusters, and the use of tools such as Kibana, Flume, and Sqoop. Additionally, the exam evaluates practical skills in Spark programming, focusing on RDD manipulation, transformations, and actions, with comparison to newer paradigms like DataFrames. Lastly, it includes a question section on machine learning basics, covering supervised and unsupervised learning, as well as distinctions between classification and regression methodologies.
This document is an exam for a Big Data module with exercises covering data storage in Hive, persistence under HDFS, troubleshooting Pig scripts, designing MapReduce workflows, and multiple-choice questions on Hadoop concepts. The exercises require understanding Hadoop's components, Pig Latin scripts, and Hive queries, alongside the mechanisms for data replication and system efficiency improvements introduced by YARN. It emphasizes practical implementation, conceptual understanding, and ecosystem utilities.
This document is an academic exam covering foundational and advanced knowledge in big data technologies, including HDFS, Hive, Spark, Flume, and Sqoop. It evaluates students' understanding of file operations in HDFS, database management with Hive, and the integration of data pipelines using Apache tools. The questions assess students' technical skills in command usage, system design, resource allocation, and code analysis. The exam emphasizes distributed computing concepts, data storage, and real-life use cases of big data frameworks.
The exam assesses foundational knowledge of Big Data, focusing on Hadoop and related technologies. It covers concepts such as Name Node failures, relational database limitations, and NoSQL database types. The assessment also includes HDFS replication advantages, CAP theorem properties, and Big Data's core attributes of volume, variety, and velocity. Lastly, it evaluates understanding of distributed file systems, and specific tools like Pig and Hive.
The document provides detailed responses to an old exam focusing on discriminant analysis and decision trees. It includes a study of company data categorized as healthy or failing. The analysis includes calculation of discriminant functions and predictions based on financial ratios.





















