Intelligence artificielle et données
307 documents à télécharger gratuitement
Cours, examens, TD, TP et exercices de intelligence artificielle et données. Thèmes couverts : machine learning, apprentissage automatique, big data, data science, fouille de données.
Ce document aborde les arbres de décision dans le cadre de l'apprentissage machine, en expliquant leurs motivations, définitions et exemples. Il décrit également l'implémentation des arbres de décision et leurs extensions, telles que le bagging et les forêts aléatoires. Les arbres de décision sont utilisés pour la classification et la régression dans l'exploration de données.
Apache Spark is a powerful open-source framework designed for efficient and unified processing of big data. It dramatically enhances performance, executing applications on Hadoop clusters up to 100 times faster in memory and supports SQL, data streaming, machine learning, and graph processing capabilities. Designed around Resilient Distributed Datasets (RDDs), it optimizes iterative computations and leverages lazy evaluation for workflow efficiency. Additionally, Spark's ecosystem includes specialized libraries like Spark Streaming, Spark SQL, MLlib, and GraphX for advanced data workflows.
This document discusses the implementation and evaluation of linear regression models using Python and the Scikit-learn library. The methodology covers key processes such as minimizing the residual sum of squares (RSS) to optimize the model's efficiency, calculating critical metrics including variance, covariance, and parameters α and β, and exploring methods to avoid overfitting. Findings also include the use of R-squared to validate model predictions, yielding a determination coefficient of 0.66.
This document introduces linear regression as a fundamental machine learning model, applied to predict the relationship between pizza size and its price using Python libraries such as numpy and matplotlib for data generation and visualization. Scikit-learn's LinearRegression model is employed for both training and prediction, demonstrating a linear relationship validated through plotted data and model equations. The methodology highlights step-by-step processes of creating data arrays, visualizing trends with graphs, and applying fit-predict functions for deriving the regression equation.
This document is an accelerated Python tutorial designed as part of the Master in Business Analytics & Data Science program. It introduces foundational Python concepts, including virtual environments, libraries, and syntax essentials. Core sections cover Python data structures, conditional logic, loops, and functions, alongside advanced topics like sorting and list comprehensions. The tutorial emphasizes Python's capabilities for data manipulation and machine learning with Scikit Learn.
Ceci est un tutorial accéléré de Python afin de voir les bases nécessaires pour le cours d’apprentissage avec Python et scikit-learn. Il couvre l'environnement Python, l'utilisation de bibliothèques telles que NumPy et Pandas, et les bases du langage. Ce document est essentiel pour les étudiants en Business Analytics et Data Science.
This document provides an overview of data interrogation languages used in Big Data environments, with a focus on PIG and HIVE. It aims to equip readers with the knowledge necessary to process and analyze large datasets effectively using these tools. The methodologies discussed include structured querying through HIVE and scripting approaches via PIG. The document emphasizes practical applications in data analytics using Hadoop ecosystems.
This document is an educational resource on Big Data, focusing on HDFS and MapReduce frameworks. It explains HDFS architecture, emphasizing data replication and fault tolerance mechanisms in a master/slave setup with NameNode and DataNode roles. Similarly, it explores the MapReduce processing model, detailing the task distribution and fault tolerance strategies. Both methods prioritize efficient, redundant, and reliable data storage and processing distributed across large clusters.
The document provides an overview of Big Data, explaining its core characteristics (4Vs: Volume, Velocity, Variety, Veracity) and discussing its applications and challenges. It introduces Hadoop as an open-source framework optimized for distributed computing, enabling storage and processing of massive datasets. Key technologies such as MapReduce, HDFS, and Hadoop's ecosystem are detailed, emphasizing cost-efficiency and scalability. Challenges like system heterogeneity, fault tolerance, and concurrency in distributed computing are identified, with potential solutions outlined.
This document delves into the implementation of OLAP for analyzing shoe sales data at "Au bon pied." A star-schema relational data structure is developed, enabling multidimensional reporting on dimensions such as Time, Model, and Store while calculating measures like sales volume and revenue. Advanced hierarchical configurations, on-the-fly calculations, and temporal trend evaluations allow for efficient analysis and decision-making. Sparse data optimization strategies further enhance storage and processing of high-dimensional sales data across thousands of references and stores.
This document introduces Online Analytical Processing (OLAP), contrasting it with Online Transaction Processing (OLTP). It details key OLAP operations, such as roll-up, drill-down, slice, and dice, as well as the underlying architecture, including data storage methods (ROLAP, MOLAP, HOLAP). Additionally, various OLAP analysis examples highlight the methodology in sales data interpretation using multi-dimensional schemas and SQL queries. Lastly, it covers challenges of physical storage, sparsity issues, and operational restructuring of cubes for better analytical insights.
This course module covers the fundamental concepts of Online Analytical Processing (OLAP) including its architecture, comparison with Online Transaction Processing (OLTP), and practical data analysis examples. Students will learn how to exploit data within a data warehouse for various analytical purposes such as summarization, consolidation, and statistical application.
This document discusses two innovative algorithms for image binarization aimed at enhancing the representation of dark objects set against clear backgrounds. The first algorithm utilizes multifrequency analysis to determine pixel classification, while the second employs local threshold learning for improved accuracy. The proposed methods address common challenges in traditional binarization techniques due to sensor and illumination inconsistencies.
This document provides an in-depth analysis of Apache Spark, a unified analytics engine for big data processing. Spark overcomes limitations of MapReduce such as high I/O latency and inefficiency in multi-pass computations. With its in-memory data processing and lazy evaluation, it achieves up to 100x faster performance in memory and up to 10x on disk compared to Hadoop. Spark allows for batch and streaming data processing, supports RDDs, DataFrames, and Datasets, and includes additional libraries for Machine Learning, graph processing, and SQL queries. Its ecosystem is expansive, with mult...
Apache Pig is a high-level platform that simplifies MapReduce programming for Hadoop using a scripting language called Pig Latin. Pig Latin provides an abstraction over Java MapReduce and allows procedural transformations with features like lazy evaluation, ETL, and user-defined functions (UDFs). The platform executes distributed tasks through directed acyclic graph (DAG) pipelines and offers integration with various storage formats. Compared to SQL, Pig caters to procedural programmers, providing flexibility, transparency, and reduced overhead in transforming and processing vast datasets.
This document introduces MapReduce as a paradigm for distributed computing, focusing on its practical advantages over traditional methods for processing large datasets. It explains the methodology through key concepts such as the map and reduce operations, the division of tasks into smaller fragments, and their parallel execution across machines in a distributed cluster. Several concrete examples, such as word frequency analysis and shared-friend calculation in social networks, illustrate the practical utility of MapReduce. These examples, coupled with systematic pseudo-codes, demonstrate h...
Hadoop is an open-source framework, developed by Doug Cutting in 2004 under Apache, for distributed processing of massive datasets (petabyte scale) using clusters of commodity hardware. It facilitates distributed storage (HDFS) and computation (MapReduce) while being fault-tolerant, scalable, economical, and performance-efficient. Its ecosystem includes tools like Hive for SQL-like querying, Pig for scripting, and Mahout for machine learning. Hadoop ensures data reliability with HDFS by partitioning data into large blocks, replicating them across nodes, and utilizing NameNodes for metadata...
The document provides a step-by-step guide for setting up the Microsoft Business Intelligence environment. It highlights the installation of Microsoft SQL Server 2014 Enterprise Edition, emphasizing compatibility and enabling key components like SSIS, SSAS, and SSRS. Additionally, it instructs the installation of SQL Server Data Tools compatible with the SQL Server version installed. The document includes links to video tutorials for further guidance.
The document outlines projects for the Master's program in Data Science. It includes a continuous assessment and a supervised assignment focused on distributed systems, specifically blockchain technology and mobile device management. A project involving the development of an online voting application is described, emphasizing voter authentication and vote confidentiality.
The document provides an in-depth overview of Big Data, its definition, importance, characteristics (5Vs: Volume, Variety, Velocity, Veracity, Value), and applications spanning industries like healthcare, marketing, politics, sports, and public security. It discusses technological frameworks like Hadoop and Spark, the methodologies for data processing, and strategies for handling limitations of traditional systems. The document emphasizes distributed systems, parallel processing, scalability, and cost-efficiency as solutions to manage and analyze extensive datasets effectively.
Ce cours vise à introduire les étudiants aux systèmes d’information décisionnels et à maîtriser leur mise en place. Il est structuré autour de quatre modules couvrant les concepts fondamentaux jusqu'à l'analyse des données. La formation se déroule à distance avec un soutien tutoriel et des séances synchrones.
The Business Intelligence (BI) program aims to train graduates capable of designing, configuring, and deploying decision support systems and knowledge management systems. It focuses on equipping learners with data analysis skills to assist decision-makers in companies. The curriculum includes a series of technical and soft skills courses over six semesters.
The report outlines an internship experience at IDsoft, a software company specializing in real estate applications. The author's main project involved contributing to the development and enhancement of the reporting capabilities of IDsoft's SaaS-based ERP solution, IDimmo. The project incorporated Business Intelligence (BI) tools such as SQL Server Integration Services (SSIS), Analysis Services (SSAS), and Reporting Services (SSRS), which were utilized for efficient data handling and reporting. The work facilitated improved data aggregation, visualization, and decision-making processes for...
This Master’s thesis focuses on designing and implementing a maintenance dashboard for CLC-Délice, using Business Intelligence (BI) principles. The methodology involves analyzing the current system, proposing improvements, and creating a dashboard to streamline maintenance management. The solution integrates multidimensional data models, ETL processes, and OLAP analysis to enhance data access and operational efficiency. The study evaluates the deployed application, showcasing significant advancements in maintenance reporting, decision-making, and cost optimization.






















