Blog (mostly math)

BITS-2 Data Visualization

Link to Coursera - BITS BSc Computer Science program: Link.

Note that the individual courses in this program are available under Coursera +.

Course - 2 : Data Visualization

Link to the course: Link.

All rights of the content go to BITS Pilani.

Refs:

ROUGH NOTES (!)
Updated: 7/8/2026

[Introducing Data Analysis]

Data Analysis: Unlocking Insights Through Data

What is Data Analysis?

The process of inspecting, cleaning, transforming, and modelling data to discover useful information.

Purpose:

  • To support decision-making

  • Find patterns

  • Gain insights

Eg:

  • Analyzing sales data to identify trends

  • Studying user behaviour on a website

  • Understanding customer feedback

Importance of Data Analysis

Why is Data Analysis important?

  • Helps in making informed decisions

  • Optimises processes and operations

  • Uncovers trends and patterns that drive strategic actions

Eg: Companies like Amazon, Netflix, and Google use data analysis to personalize user experiences.

Types of Data

Structured Data (Organized data)

Eg: Sales records, student grades, etc. Usual formats: Databases and spreadsheets.

Unstructured Data (Unorganized data)

Eg: Emails, social media posts, etc. Usual formats: Text, Image, Video.

Semi-structured Data (Partially organized data)

Eg: Log files, metadata, etc. Usual formats: Json, xml.

Types of Data Analysis

Descriptive Analysis:

Summarises historical data.

Eg: Average sales per month.

Diagnostic Analysis:

Investigates reasons behind past outconmes.

Eg: Why did sales drop last quarter?

Predictive Analytics:

Uses historical data to forecast future outcomes.

Eg: Projecting next quarter’s sales.

Prescriptive Analytics:

Suggests actions based on predictions.

Eg: Recommending marketing strategies to boost sales.

Common Techniques in Data Analysis

Statistical Analysis: Mean, Median, Mode, Standard Deviation, etc.

Data Visualization: Charts, Graphs, Heatmaps, and more to visually represent data.

Correlation and Regression Analysis: Understanding relationships between variables.

Machine Learning Techniques: Clustering, Classification, and Prediction Models.

Tools for Data Analysis

Spreadsheet tools (eg: Excel, Google Sheets)

  • Basic analysis, pivot tables, charts

Statistical tools (eg: R, Python with Pandas, SPSS)

  • Advanced statistical analysis and modelling

Data Visualization Tools (eg: Tableau, Power BI)

  • Creating interactive charts and dashboards

Database Management Systems (eg: SQL)

  • Querying and managing large datasets

Getting Started with Data Analysis

Start Small: Begin with simple datasets and basic tools like Excel then proceed to visualisation tools or programming libraries.

Practice: Work on real-world projects or datasets.

Learn Continuously.

[Types of Data Analysis]

Types of Data Analysis

  • Descriptive Analysis

  • Diagnostic Analysis

  • Predictive Analysis

  • Prescriptive Analysis

  • Exploratory Analysis

Descriptive Analysis

Describes and summarises past data to identify patterns, trends, and distributions.

Key techniques:

Statistical Summaries: Mean, median, mode, standard deviation

Visualisation Tools: Bar charts, histograms, pie charts, frequency distributions

Descriptive Statistics: Measures of spread (range, quartiles), shape (skewness, kurtosis)

Descriptive Analysis: Applications

Business: Analysing monthly sales data to determine which products performed well.

Healthcare: Summarising patient demographics to understand disease distribution.

Social Media: Evaluating engagement metrics like likes, shares and comments.

Eg: A retailer summarises quarterly sales data across different regions to identify top-performing locations.

Diagnostic Analysis

Delves into the data to discover the causes behind trends or anomalies.

Key techniques:

Data mining: Discovering patterns within large datasets

Root Cause Analysis: Identifying underlying reasons for observed results.

Correlation and Regression Analysis: Exploring relationships between variables.

Diagnostic Analysis: Applications

Healthcare: Investigating the causes of a spike in hospital admissions.

Marketing: Analysing reasons behind a sudden drop in website traffic.

Manufacturing: Diagnosing the causes of product defects.

Eg: An e-commerce platform analyses customer churn data to find out why users are abandoning their carts.

Predictive Analysis

Uses historical data and statistical models to predict future outcomes.

Key techniques:

Machine Learning: Algorithms like regression, decision trees, neural networks

Time Series Analysis: Forecasting trends based on time-indexed data

Predictive Modelling: Using models to simulate possible outcomes (eg, Monte Carlo simulations)

Predictive Analysis: Applications

Finance: Predicting stock prices or market trends.

Retail: Forecasting product demand for the upcoming season.

Healthcare: Predicting disease outbreaks or patient readmissions.

Eg: A financial institution uses predictive modelling to estimate credit risk and default probabilities.

Eg: An online streaming service predicts user preferences to recommend content.

Prescriptive Analysis

Provides recommendations on actions to take, based on predictive insights, to optimise outcomes.

(What happened -> What will happen -> How can we make it happen)

Key techniques:

Optimisation Models: Linear programming, optimisation algorithms to find the best course of action.

Simulation: Testing various scenarios to evaluate potential outcomes.

Decision Analysis: Tools like decision trees, and cost-benefit analysis.

Prescriptive Analytics: Applications

Supply Chain: Optimising inventory levels based on demand forecasts.

Healthcare: Personalised treatment plans based on patient data.

Marketing: Designing targeted advertising campaigns.

Eg: A logistics company uses prescriptive analytics to optimise delivery routes, reducing costs and time.

Eg: A retail chain adjusts pricing strategies based on predictive insights to maximise profits.

Exploratory Analysis

Uncover patterns, relationships, or insights in data without having a specific hypothesis in mind.

Key techniques:

Data Visualisation: Scatter plots, heatmaps, box plots to reveal trends and correlations.

Clustering Algorithms: K-means clustering to group similar data points.

Dimensionality Reduction: Techniques like PCA (Principal Component Analysis) to simplify data.

Exploratory Analysis: Applications

Product Development: Exploring user feedback to discover unmet needs.

Healthcare: Uncovering patterns in patient data that may suggest new research directions.

Social Media: Identifying emerging trends or viral content.

Eg: A company explores customer review data to identify common complaints and potential areas for product improvement.

[Comparing Types of Data Analysis]

Descriptive vs Diagnostic

What happened? vs Why did it happen?

Summarises historical data vs Investigates underlying causes.

Eg: A company notices a drop in sales (descriptive) and then analyses customer feedback to understand why (diagnostic).

Predictive vs Prescriptive

What might happen? vs What should we do about it?

Forecasts future outcomes based on historical data vs Recommends actions to optimise outcomes.

Eg: A retailer predicts future demand (predictive) and then adjusts inventory levels accordingly (prescriptive).

Choosing the Right Type of Analysis

Factors to consider:

Objective

  • What is the goal of the analysis?

  • Understanding the past, predicting the future, or making decisions.

Data Type

  • Structured vs unstructured data

Resources

  • Tools, time, and expertise available

Factor 1: Objective

Goal of the Analysis: Determine what you want to achieve.

  • Understanding the past: Use descriptive analysis to summarise historical data.

  • Predicting the future: Use predictive analysis to forecast future trends and outcomes.

  • Making decisions: Use prescriptive analysis to recommend actions based on predictions.

  • Discovering insights: Use exploratory analysis to uncover new patterns and insights.

Factor 2: Data Type

Structured Data:

  • Organised in rows and columns (eg: databases, spreadsheets)

  • Suitable for: Descriptive, Diagnostic, Predictive, and Prescriptive analysis.

Unstructured Data:

  • Raw data not organised in a predefined manner (eg: text, images)

  • Suitable for: Exploratory analysis, often requiring data preprocessing.

Factor 3: Resources

Tools:

  • Availability of software and tools (eg: Excel, R, Python, BI Tools)

Time:

  • Time constraints for analysis and decision-making.

Expertise:

  • Skills and knowledge required for different types of analysis (eg: statistical expertise, machine learning knowledge)

[Exploratory vs Explanatory Data Analysis]

Introduction to Data Analysis

Data analysis is the process of examining, transforming, and modelling data to discover useful information, inform conclusions, and support decision-making.

Purpose: Understanding data patterns, relationships, and trends.

Types:

  • Exploratory Data Analysis (EDA)

  • Explanatory Data Analysis

Exploratory Data Analysis (EDA)

EDA is an approach to analysing data sets to summarise their main characteristics, often with visual methods.

Purpose:

  • Discover patterns

  • Spot anomalies

  • Test hypotheses

  • Check assumptions

Techniques:

  • Descriptive statistics (mean, median, mode)

  • Data visualisation (histograms, box plots, scatter plots)

  • Data cleaning and transformation

Population -> Data (some observed portion of the population) -> Inference

Why is EDA important?

Understanding Data: Before diving into complex modelling, its essential to understand what the data looks like.

Data Quality Check: Identify missing values, outliers, and inconsistencies.

Hypothesis Generation: Develop initial hypotheses based on data patterns.

Eg: Exploring sales data to find trends, seasonality, and anomalies before building predicitve models.

What is Explanatory Data Analysis?

Explanatory Data Analysis focuses on explaining or presenting the results of an analysis to a specific audience in a clear and concise manner.

Purpose:

  • Communicate Insights

  • Validate or refuse hypotheses

  • Support decision-making

Techniques:

  • Focused visualisations (bar charts, pie charts, line graphs)

  • Storytelling with data

  • Simplification and emphasis on key findings

Why Explanatory Data Analysis is Important?

Communication: Translate complex data insights into understandable information for stakeholders.

Decision Support: Helps in making informed decisions based on the analysis.

Persuasion: Effective storytelling can influence opinions and actions.

Eg: Presenting quarterly financial data to management to support business decisions.

EDA vs Explanatory Analysis

Objective: “Discovery and exploration” vs “Communication and presentation”

Audience: “Analysts and data scientists” vs “Stakeholders, decision-makers, general audience”

Tools and Techniques: “Interactive visualisations, complex charts” vs “Simple, focused visualisations, clear narratives”

EG: “Involves diving deep into data with scatter plots and correlation matrices” vs “Involves summarising key findings with a few well-chosen charts”

EDA Tools:

  • Python (Pandas, Seaborn, Matplotlib)

  • R (ggplot2, dplyr)

  • Jupyter Notebooks

Explanatory Tools:

  • Tableau

  • Google Data Studio

  • PowerPoint

  • Excel

Tldr: EDA is about discovering insights, and Explanatory Data Analysis is about communicating those insights to your audiences.

[Data Analysis Process]

Five key steps:

  • Define the Problem or Objective

  • Collect the Data

  • Clean the Data

  • Analyse the Data

  • Interpret and Communicate Results

Iterative Nature: Data analysis is often cyclical, requiring refinement of earlier steps based on insights gained in later steps.

Define the Problem or Objective

Why is this step important?

  • Understanding the context and objective helps guide the entire analysis.

Key activities:

  • Identify the research question or business problem.

  • Define success criteria and the scope of the analysis.

  • Set hypotheses (if applicable)

Collect the Data

Types of Data Collection

  • Primary Data: Surveys, Experiments, Observations.

  • Secondary Data: Public Datasets, Company Databases, Previously Collected Data

Clean the Data

(This takes up about 60 to 70 % of your time)

Why is Data Cleaning critical?

  • Garbage in, garbage out - the quality of your analysis depends on clean data.

Key Tasks:

  • Handling missing data (imputation i.e. estimating missing information, deletion)

  • Correcting errors and inconsistencies

  • Removing duplicates and irrelevant information

  • Standardising formats (eg: dates, units)

Analyse the Data

EDA:

  • Understand the structure, patterns, and relationships in your data

  • Techniques: Descriptive Statistics, and visualisations (eg: histograms, scatter plots)

Advanced Analysis:

  • Statistical Analysis, Hypothesis testing, regression, ANOVA

  • Machine Learning: Classification, clustering, predictive modelling

Tools can be Excel, R, Python (pandas, NumPy, scikit-learn), SQL.

Interpret and Communicate Results

Analysis is only valuable if others can understand and act on it.

Key elements of Interpretation:

  • Drawing meaningful conclusions

  • Linking findingd back to the original objectives

  • Identifying limitations and assumptions

Communication Formats:

  • Reports: Structured and detailed.

  • Presentations: Engaging and focused on key insights.

  • Visualisations: Charts, graphs, dashboards (eg: Tableau, Power BI)

A Data Analysis Case Study

Scenario: Analysing customer churn (i.e customers leaving) in a subscription-based service.

Defining the Problem: Reducing churn rate

Collecting Data: Customer usage, demographics, feedback

Cleaning Data: Removing duplicates, handling missing values

Analysing Data: Identifying patterns, Applying predictive models

Interpreting Results: Actionable insights for retention strategies

[Challenges in Data Analysis]

Definition of Data Analysis: The process of inspecting, cleansing, transforming, and modelling data to discover useful information, draw conclusions, and support decision making.

Data Quality Issues

Missing Data:

Causes and impact on analysis

Strategies for handling missing values (eg: imputation, deletion)

Outliers and Noise:

Affect data integrity

Methods to identify and manage them

Data Inconsistency:

Examples of inconsistencies

Techniques to ensure data uniformity

Data Integration Challenges

Data Sources:

  • Problems with combining data from multiple sources (eg: databases, APIs)

Schema Mismatch:

  • Issues with different data formats and structures

  • Solutions like data transformation and normalisation

Data Silos:

  • Isolated data storage systems impact analysis and strategies for integration

Complexity of Data

High Dimensionality:

Challenges with large datasets and methods like dimensionality reduction (eg: PCA)

Unstructured data:

Difficulties in analysing text, images, and other unstructured data types.

Tools and techniques for handling unstructured data.

Data Privacy and Security

Data Protection Laws: Regulations like GDPR and CCPA

Ethical Considerations: Ensuring responsible use of data and avoiding misuse

Security Sensitive Data: Techniques for encrypting and protecting

Scalability Issues

Handling Large Datasets: Challenges in processing and analysing big data

Performance Optimisation: Techniques for improving the efficiency of data analysis (eg: distributed computing, parallel processing)

Human Factors

Cognitive Biases: Biases affect data interpretation and decision-making

Skill Gaps: Challenges related to the expertise of data analysts and the need for ongoing training

Tools and Technology

Software Limitations: Issues with available tools and software for data analysis.

Keeping Up with Technological Advances: The need for staying updated with the latest tools and technologies.

[Data Visualisations]

Data visualisation (in the form of dashboards)

  • The goal of making data more accessible and understandable

  • The go-to tool for many businesses to analyse and share information

What is Data Visualisation?

Data Visualisation is the graphical representation of data and information.

Data Visualisation tools provide an accessible way:

  • to see and understand trends, outliers, and patterns in data

  • Using visual elements like charts, graphs, and maps

Provides an excellent way for employees or business owners to present data to non-technical audiences without confusion.

Advantages and Disadvantages

Presenting data graphically may appear straightforward and beneficial.

Misrepresentation or misinterpretation can occur if the wrong visualisation style is used.

Crucial to consider both advantages and disadvantages when creating data visualisations.

Advantages

Our eyes are drawn to colors and patterns, making visual distinctions easy.

  • Visual culture includes art, ads, TV, and movies, and data visualisation fits within this context.

Captures attention and effectively conveys messages.

Quickly identify trends and outliers, aiding faster internalisation of information.

A purposeful form of storytelling.

Disadvantages

Disadvantages of data visualisation may not always be obvious.

Visualisations with many data points can lead to inaccurate assumptions.

  • Poorly designed visualisations can be biased or confusing.

  • Potential for biased or inaccurate information.

  • Misinterpretation of correlation as causation.

  • Core messages can get lost in translation.

Why is Data Visualisation Important?

Data visualisation helps people see, interact with, and better understand data.

  • Effective visualisations align everyone, regardless of expertise level.

  • Most professional industries benefit from making data more accessible.

Practical applications of data visualisation are widespread and valuable

  • Is essential in STEM fields, government, finance, marketing, history, consumer goods, services, education, sports, and more.

Modern professionals benefit from combining creative storytelling with technical analysis.

Data visualisation bridges the gap between analysis and visual storytelling.

[Why Visual Analytics?]

Visual analytics combines the power of data visualisation with advanced analytical techniques to uncover deeper insights.

Purpose: Enable users to explore complex data interactively, allowing them to ask and answer ‘why’ questions in real-time.

Combines the strengths of traditional visualisation methods with more sophisticated analytical techniques, bridging the gap between raw data and actionable insights.

Importance of asking ‘Why’

Curiosity in Data Exploration

  • Curiosity drives deeper understanding and discovery in data analysis.

Explanation:

  • Encourages analysts to go beyond surface-level insights and question the reasons behind observed patterns and anomalies.

Eg:

Investigating unexpected spikes in sales data by asking why they occured can reveal underlying factors such as seasonal trends, marketing impacts, or changes in consumer behaviour.

Visual Analytics Advantage

  • Visual analytics empowers users to dynamically question and explore data

Explanation:

Provides interactive tools and visualisations that facilitate questioning and discovery, making it easier to uncover underlying causes and trends.

Eg:

Interactive dashboards allow users to drill down into specific data points, apply filters, and adjust parameters to see how changes impact the overall picture.

Benefits of Visual Analytics

Enhanced Decision-Making

  • Visual analytics facilitates informed and timely decision-making

How?

  • By presenting data in a clear, intuitive format, it highlights actionable insights and supports strategic planning.

Eg:

Decision makers can quickly assess the impact of different business strategies through interactive dashboards and visualisations, leading to more data-driven choices.

Spotting Trends: Visual analytics enables rapid identification of trends and anomalies.

Visualisation tools can reveal patterns that might be missed in raw data, such as emerging market trends or unusual deviations.

Eg:

Line charts and heat maps can highlight seasonal sales trends or unexpected spikes in website traffic.

Improving Data Literacy: Visual analytics enhances the ability of non-technical users to understand and work with data

Simplifies complex data and makes it accessible to individuals without a technical background, promoting better data-driven decision-making across the organisation.

Eg:

Interactive visualisations with user-friendly interfaces help employees from various departments interpret data without needing advanced statistical knowledge.

Use Case: Business Intelligence

Scenario: A retail company uses visual analytics to analyse sales data across regions

Impact: The company identifies underperforming areas and optimises inventory, leading to increased sales.

[Visual Analysis for Everyone]

Challenges of Traditional Data Analysis

1 Complexity

Data Overload:

Traditional analysis methods often involve complex data sets and require specialised knowledge to interpret effectively

Analytical Tools:

Advanced tools and techniques used in traditional analysis can be difficult to master and require significant training

Data Preparation:

Extensive data cleaning and preparation are often needed before meaningful analysis can occur

2 Accesibility

Technical Barriers:

Non-technical users may struggle with traditional data analysis tools and methods, limiting their ability to independently access and understand data insights

Dependence on Experts:

Accessing data insights frequently depends on data analysts or IT professionals, creating bottlenecks and reducing data democratisation

Complex Interfaces:

Many traditional tools have complex interfaces that can be intimidating for casual users or those without a technical background

3 Speed

Slow Insight Generation:

Traditional methods can lead to delays in generating insights due to lengthy data processing and analysis cycles

Manual Processes:

Manual data handling and report generation contribute to longer turnaround times for actionable insights

Inflexibility:

Adapting to new questions or changing data requires significant rework, slowing down the responsiveness to emerging business trends

Introduction to Visual Analysis

Visual analysis involves using graphical representation of data to uncover patterns, trends, and insights that might be obscured in raw data formats.

Combines data visualisation with analytical techniques to facilitate a deeper understanding of information.

Significance in Today’s Data-Driven World

Increasing Data Volume: With the exponential growth of data, visual analysis helps to make sense of large volumes of information quickly and effectively.

Decision Making: It enhances decision-making by providing clear and actionable insights through intuitive and interactive visualisations.

Accessibility: Makes data insights accessible to a wider audience, including those without a technical background.

How Visual Analysis Democratises Data Insights?

Empowerment for All Users: Visual analysis tools enable users at all levels, from executives to operational staff, to interact with and understand data without needing advanced analytical skills.

Enhanced Data Literacy: By simplifying complex data, visual analysis fosters a culture of data literacy and encourages more informed decision-making across the organisation.

Collaborative Insights: Encourages collaboration by allowing teams to share interactive visualisations and insights, leading to better alignment and more strategic planning.

Implementing Visual Analysis

1 Introduce Visual Analysis Tools

Selection:

Choose the right visual analysis tools (eg, Tableau) that align with your organisation’s needs and data infrastructure.

Deloyment:

Implement the tools across the organisation, ensuring they are integrated with existing data sources and systems.

2 Train Users

Training Programs:

Develop and deliver training sessions to equip users with the skills needed to effectively use visual analysis tools.

Hands-On Workshops:

Offer practical workshops where users can work with real data and learn how to create and interpret visualisations.

3 Establish Best Practices

Standardisation:

Develop and document standard procedures for creating and sharing visualisations to ensure consistency and quality.

Support:

Set up support channels to assist users with technical issues and best practices for visual analysis.

[The Future of Visual Analytics]

Why the Future of Visual Analytics Matters?

Two Reasons:

  • The growing importance of data-driven decision making

  • Technological advancements shaping the field

The growing importance of data-driven decision making

Informed decisions:

As organisations increasingly rely on data for strategic decisions, visual analytics becomes crucial for transforming raw data into actionble insights.

Clear, effective visualisations allow decision-makers to quickly grasp complex information, leading to more accurate and timely decisions.

Competitive advantage:

Companies that excel in data-driven decision-making outperform their peers.

Visual analytics enables businesses to identify trends, forecast outcomes, and respond proactively to market changes, thereby gaining a competitive edge.

Wider Accessibility:

Data-driven decision making is no longer limited to data scientists or analysts.

With the rise of user-friendly visualisation tools, professionals across various domains can leverage data insights, democratising the use of data throughout organisations.

Adaptation to Complex Data Environments:

As data sources grow in variety and volume (eg: IoT, social media, sensor data), traditional data analysis methods struggle to keep up.

Visual analytics offers an intuitive way to interpret complex datasets, making sense of large-scale data more effectively.

Technological advancements shaping the field

AI and Machine Learning

  • The integration of AI and machine learning into visual analytics is transforming how data is analysed.

  • These technologies enable automated pattern recognition, predictive insights, and adaptive visualisations, which can adjust in real-time as new data flows in.

Augmented and Virtual Reality

  • Immersive technologies are pushing the boundaries of how data is visualised.

  • AR and VR (Virtual Reality VR immerses users in digital environments, and Augmented Reality AR overlays digital elements onto the real world) offer new ways to explore multi-dimensional datasets, allowing users to interact with data in a more intuitive and engaging manner.

Cloud Computing:

Cloud-based visual analytics platforms provide scalability, anabling organisations to process and visualise vast amounts of data without being constrained by local infrastructure.

This shift also supports real-time collaboration and access to data from anywhere, driving more agile decision making.

Natural Language Processing (NLP):

NLP advancements are making data querying more accessible.

Users can now interact with visual analytics platforms using natural language, asking questions and recieving visual insights without needing advanced technical skills.

Real-Time Data Processing:

The ability to visualise data in real-time is becoming increasingly important in industries such as finance, healthcare, and logistics.

Technological advancements in data streaming and processing enable the creation of dashboards and visualisations that update instantly, allowing for immediate responses to emerging trends or issues.

Future Directions and Predictions

Convergence of Visual and Predictive Analytics:

  • Integration of forecasting models within visual dashboards

  • Use of simulations in visual analytics

Automation and No-Code Solutions:

  • Rise of no-code/low-code platforms for data visualisation

  • Implications for democratising visual analytics

Interdisciplinary Collaboration:

  • Cross-industry applications and collaboration

  • The importance of working with domain experts

[Distinguishing Data Roles]

Data Roles

Data plays a crucial role in decision-making, innovation, and strategic planning.

Various specialised roles exist to manage and analyse data effectively.

Key Roles

Data Analyst

Data Scientist

Data Engineer

Data Architect

Database Administrator

Data Steward

Business Intelligence (BI) Analyst

Data Analyst

Responsibilities:

  • Analyse and interpret complex data sets

  • Create reports and dashboards

  • Provide actionable insights to support business decisions

Skills:

Data visualisation tools (eg Tableau, Power BI)

Statistical analysis

Excel, SQL, R, Python

Data Scientist

Responsibilities:

Design and develop machine learning models

Perform advanced statistical analyses

Predict trends and behaviours using data

Skills:

Python, R, SQL

Machine learning frameworks (eg: Tensorflow, PyTorch)

Data wrangling and preprocessing

Data Engineer

Responsibilities:

Build and maintain data pipelines

Ensure data is available, clean and well-structured

Collaborate with data scientists and analysts to optimise data flow

Skills:

Programming (Python, Java, Scala)

Data storage systems (eg Hadoop, Spark)

Cloud services (AWS, Azure, Google Cloud)

Data Architect

Responsibilities:

Design and structure databases and data systems

Ensure data architecture aligns with business goals

Oversee data integration and management strategies

Skills:

Database design and modelling

ETL (Extract, Transform, Load) processes

Knowledge of various data storage solutions

Database Administrator

Responsibilities:

Manage and maintain databases

Ensure data integrity and security

Optimize database performance

Skills:

Database management systems (eg: MySQL, Oracle, SQL Server)

Backup and recovery techniques

Security protocols and compliance

Data Steward

Responsibilities:

Ensure data quality and governance

Establish data policies and standards

Monitor data usage and compliance

Skills:

Data governance and frameworks

Communication and collaboration

Knowledge of regulatory requirements (eg GDPR, HIPAA)

Business Intelligence (BI) Analyst

Responsibilities:

Develop BI strategies to support business goals

Design and create data visualisations and dashboards

Work closely with stakeholders to identify reporting needs

Skills:

BI tools (eg Tableau, Power BI)

SQL and data warehousing

Analytical thinking and business acumen

Collaboration between Roles

Importance of collaboration between different data roles:

Data Engineers ensure data availability for Analysts and Scientists.

Data Scientists build models that are deployed by ML Engineers.

Data Architects ensure the infrastructure supports all roles.

BI Analysts translate data insights into actionable business strategies.

[Introducing Data Visualisation Analyst (Visual Analyst)]

Who is a Data Visualisation Analyst?

A Data Visualisation Analyst is a professional who creates visual representations of data to help organisations understand complex information.

Purpose: To convert raw data into clear, compelling visuals that facilitate decision-making and insight discovery.

Role in Data Analytics

Bridges the Gap

Between Raw Data and Insights:

  • Transforms complex datasets into understandable visual formats

  • Purpose: Makes data accessible and actionable for stakeholders

Effective Visual Communication:

  • Techniques: Utilizes charts, graphs, dashboards, and interactive reports

  • Goal: Communicate findings clearly and support data-driven decisions

Responsibilities of a Data Visualisation Analyst

Creating Visual Representations of Data:

Transform raw data into clear, insightful visuals

Designing Dashboards:

Provide a comprehensive view of key metrics and data points

Designing Interactive Reports:

Allow users to explore data dynamically and engage with visual elements

Skills required for a Data Visualisation Analyst

Proficiency in Visualisation Tools:

Tableau, Power BI, Google Data Studio

Skill Areas:

Data import/export, chart customisation, dashboard creation

Understanding of Design Principles:

Contrast, Alignment, Proximity, Repetition

Storytelling with Data

Elements of Storytelling: Crafting a narrative, focusing on the key message, structuring the visual flow

Techniques: Using annotations, highlighting trends, guiding the viewer through the data

Becoming a Data Visualisation Analyst

Building a Portfolio

Showcase Skills through Projects:

  • Real-World Projects: Include examples of work done for clients or as part of professional experience

  • Case Studies: Document the problem, solution, and impact of your visualisation

  • Sample Dashboards: 1 Diverse examples: Create dashboards for various industries or data types (eg sales, finance, healthcare). 2 Interactive elements: Demonstrate your ability to design interactive and user-friendly dashboards

Professional Development

Joining Communities

Professional Associations: Join organizations such as the Data Visualisation Society or local meetup groups.

Online Communities: Participate in forums like Reddit’s r/dataisbeautiful or Linkedin groups

Attending Workshops and Conferences:

Events: Attend industry conferences, workshops, and webinars focused on data visualisation and analytics

Networking: Engage with peers, share knowledge, and learn about new tools and techniques

[Installable Tools]

What are Installable Tools?

Software that needs to be downloaded and installed on a computer

Offline functionality and data security

Customisation and integration capabilities

Advantages:

Enhanced performance and speed

Better handling of large datasets

Advanced customisation and flexibility

Offline access and data privacy

Tableau Desktop

Features:

Drag-and-Drop Interface

Rich Visualisation Options

Supports Data Blending

Use Cases:

Business Intelligence

Complex Data Analysis

Pros:

User-friendly

Powerful Analytics

Wide Range of Visualisations

Cons:

Cost

Learning Curve

Microsoft Power BI Desktop

Features:

Interactive reports

Dashboards

Integration with MS products

Use cases:

Business reporting

Data exploration

Pros:

Integration with Excel

Cost-effective

Extensive data connectivity

Cons:

Limited advanced analytics

Occasional performance issues with large data

Qlik Sense Desktop

Features:

Self-service Data Visualisation

Associative Data Model

Interactive Dashboards

Use Cases:

Data Discovery

Interactive Analytics

Pros:

Intuitive Interface

Strong Data Association Capabilities

Wide Range of Visualisations

Cons:

Complex Licensing

Limited Advanced Features in Desktop Version

IBM Cognos Analytics

Features:

AI-driven analytics

Visualisations

Report generation

Use cases:

Enterprise reporting

Predictive analytics

Pros:

Advanced analytics

Integration with other IBM products

Cons:

High cost

Steep learning curve

D3.js (Desktop Environment)

Features:

Highly customisable visualisations using JavaScript

Use Cases:

Custom data visualisations

Web integration

Pros:

Flexibility

Detailed control over visualisations

Cons:

Requires programming skills

More time-consuming setup

[Libraries and Packages]

Popular Visualisation Libraries

Python Libraries:

Matplotlib: Basic plotting and customisation

Seaborn: Statistical plots and themes

Plotly: Interactive plots and dashboards

Bokeh: Interactive web-based plots

Altair: Declarative statistical visualisation

R libraries:

ggplot2: Grammar of graphics for elegant visualisations

Shiny: Interactive web apps and dashboards

Lattice: Multi-panel data visualisations

Python Libraries - Matplotlib

Matplotlib is a basic plotting library for Python that provides a wide range of fuctionality for creating static visualisations.

One of the oldest and most widely used libraries in the Python ecosystem for data visualisation.

Use Cases:

Matplotlib is suitable for creating static visualisations where interactivity is not required

Particularly useful for producing high-quality plots for publications or reports

Its flexibility allows for extensive customisation, making it a go-to tool for a wide range of visualisation needs

Python Libraries - Seaborn

Seaborn is a powerful data visualisation library built on top of Matplotlib.

Simplifies the process of creating attractive and informative statistical plots

Offering high-level interfaces for drawing appealing and complex visualisations effortlessly

Advanced Statistical Plots: Includes built-in support for complex plots like pair plots, joint plots, and regression plots

Seaborn is ideal for exploring and visualising data distributions, especially when dealing with statistical relationships between variables

Python Libraries - Plotly

Plotly is a versatile Python library that specialises in creating interactive plots and dashboards

Allows users to generate visualisations that can be embedded in web applications, providing dynamic and responsive data exploration experiences

Features:

Interactive Elements, 3D plots, Animations

Use Case:

Plotly is ideal for creating interactive, web-based visualisations that allow users to explore data dynamically

Particularly useful for developing dashboards and visual analytics applications that require real-time interaction

Python Libraries - Bokeh

Bokeh is a powerful Python library designed for creating interactive visualisations that can be easily integrated into web applications.

Enables the development of rich, interactive plots and dashboards with a focus on web deployment

Features:

Real-Time Updates: Supports streaming and real-time data updates, making it ideal for dynamic visualisations

Custom Widgets: Provides a variety of custom widgets

Use Case:

Best suited for creating interactive visualisations within web applications

R Libraries - ggplot2

ggplot2 is a widely used R library based on the Grammar of Graphics, which provides a consistent and flexible framework for creating data visualisations.

Simplifies the process of building complex plots by allowing users to layer components on top of each other

Features:

Layered Graphics: Allows users to build visualisations by adding layers such as points, lines, bars, and more, enabling the creation of sophisticated multi-layered plots.

Use Case:

Ideal for creating complex, multi-layered visualisations in R

Especially useful for exploratory data analysis, where users need to visualise relationships, distributions, and trends across different variables

R Libraries - Shiny

Shiny is an R package that allows users to build interactive web applications directly from R

Bridges the gap between data analysis and web development, enabling the creation of powerful and dynamic dashboards without requiring extensive knowledge of HTML, CSS, or Javascript

Features:

Server-Side and Client-Side Code Integration

Interactive UI Components

Real-Time Interaction

Use Case:

Shiny is perfect for creating interactive dashboards and data driven applications, especially when users need to manipulate data and see instant visual feedback

Comparision of Libraries

[Cloud Based Platforms]

Cloud-based Data Visualisation Platforms

Cloud-based data visualisation platforms offer powerful tools for analysing and presenting data through interactive and accessible dashboards

Leverage cloud technology to provide scalable, flexible, and collaborative solutions

Some Popular Choices:

Tableau Online

Microsoft Power BI

Google Data Studio

Qlik Sense Cloud

Why Cloud-Based Platforms?

Scalability: Handle large datasets and complex visualisations

Accessibility: Access from anywhere with an internet connection

Collaboration: Real-time collaboration and sharing capabilities

Cost-Efficiency: Pay-as-you-go pricing models

Integration: Seamless integration with other cloud services

Tableau Online

Cloud-based data visualisation platform that provides robust tools for creating interactive and dynamic dashboards, offers a user-friendly experience and a broad range of visualisation options

Key Features:

Interactive dashboards: Create and interact with dynamic dashboards that update in real-time

Ease of Use: Intuitive interface with drag-and-drop functionality for easy visualisation creation

Extensive Visualisation Options: Wide variety of charts, graphs, and visual elements to represent data effectively

Microsoft Power BI

Microsoft Power BI is a cloud-based business analytics platform that provides comprehensive tools for data visualisation and reporting

Designed to integrate seamlessly with the Microsoft ecosystem, offering advanced data modelling and extensive visualisation options

Key Features:

Integration with Microsoft Ecosystem: Seamlessly connects with Microsoft products such as Excel, Azure, and SharePoint for a unified data experience

Powerful Data Modelling: Advanced capabilities for creating complex data models, including relationships, calculations and aggregations

Extensive Visualisation Tools: Wide range of visualisation options including charts, maps, and custom visuals to effectively present data insights

Google Data Studio

A free, cloud based data visualisation and reporting tool that allows users to create interactive and customisable reports

Integrates seamlessly with various Google services, making it a convenient option for data analysis and visualisation

Key features:

Integration with Google Services: Easily connect to Google Analytics, Google Sheets, Google Ads, and other Google products for a cohesive data experience.

Customisable Reports: Design and tailor reports to meet specific needs, with a variety of charts, tables, and widgets to choose from

Free to Use: No cost to access the platform, making it an economical choice for individuals and businesses

Qlik Sense Cloud

Qlik Sense Cloud is a cloud-based data visualisation and anlytics platform that emphasises an associative data model and advanced analytics capabilities

Offers a user-friendly interface designed for both data exploration and interactive dashboard creation

Key features:

Associative Data Model: Allows users to explore data freely and discover insights without predefined query paths, providing a more intuitive data interaction experience

Advanced Analytics: Supports complex data analysis, including predictive analytics and advanced calculations, to uncover deeper insights

User-Friendly Interface: Intuitive design with drag-and-drop functionality, making it accessible for users of all skill levels

Comparision of Cloud-Based Platforms

[Quantitative Relationships]

Numbers that Matter!

Numbers themselves are neither inherently boring nor interesting.

Appeal depends on the significance and relevance of the information they convey.

Quantitative information is simply data that measures the quantity of something.

The impact and appeal of any information, whether quantitative or not, stems from the importance of the message it communicates.

Why to Understand Numbers?

Numbers alone are not useful unless they measure something significant.

As a communicator, it is your responsibility to present information clearly and in an easily understandable manner.

Designing tables and graphs is exciting, crucial first to understand the numbers that form the basis of these visualisations.

Quantitative Information

The choice between using a table or a graph to display quantitative information depends on the story you want to tell.

Quantitative stories focus on relationships between data points.

Eg:

Elements of Quantitative Story

Quantitative stories involve relationships between measures of quantity and associated categories (eg geography, time)

These stories feature two data types:

  • Quantitative (measuring things)

  • Categorical (grouping information)

Categories include items like geographical areas (eg north, east, south, west) or time periods (eg months).

Use of Categories and Measures

Quantitative values and categorical items serve complementary purposes, typically structured and displayed differently.

The distinction between quantitative values (measures) and categorical items is crucial for creating effective tables and graphs.

Quantitative Relationships

Can be simple associations between quantitative values and their categorical labels.

  • Measures linked to categories like geography or time.

More complex relationships involve direct associations between multiple sets of quantitative values.

  • Marketing emails sent versus orders recieved.

  • Handwashing frequency versus infection rates in hospitals.

[Relationships Within Categories]

Categorical items used in tables and graphs to label corresponding measures can relate to one another in the following ways:

Nominal

Ordinal

Interval

Hierarchical

Nominal Relations

Values in a single category are discrete and have no intrinsic order.

Eg: Sales regions like East, West, North, and South have no inherent order.

The term “nominal” means “in name only”.

In a quantitative story that is nominal in nature:

  • Quantitative values are associated with categorical labels

  • The story does not relate the categorical items to one another in any specific way.

Ordinal Relations

In an ordinal relationship, the categorical items have a prescribed order.

Typical examples include:

  • first, second, third, …

  • small, medium, large, …

  • best salesperson, second best salesperson, …

To display them in any other order, except in reverse, would rarely be meaningful.

Interval Relations

Categorical items consist of a sequential series of numerical ranges.

These numerical ranges, called intervals, subdivide a larger range of quantitative values into smaller ranges.

Intervals are arranged in order from smallest to largest.

Eg:

To analyse the distribution of order sizes, individual order amounts are grouped into sequential, equally sized intervals.

Eg: An Interval Relation called Time-Series Relation.

Hierarchical Relationship

Previous categorical relationships focused on items within the same category.

A hierarchical relationship involves multiple categories connected in a “parent-to-child” structure

  • Organise quantitative information effectively in tables

  • Visually represent how categories are structured and connected across different levels

Eg:

[Relationships between Quantities]

Relationships between Quantities

Categorical items can also relate to one another by virtue of the quantitative values associated with them.

The quantitative values can be arranged to display the following relationships:

Ranking

Ratio

Correlation

Ranking

In ranking relationship, the categorical items are arranged based on their quantitative values, either ascending or descending.

For example:

Arranging your company’s top five sales orders of the current quarter by revenue from largest to smallest enhances the narrative and clarity of the data presentation.

Ratio

Compares two quantitative values by dividing one by the other to express their relative quantities.

Often expressed as percentages where the whole equals 100% and the part represents a lesser percentage.

Simplifies comparisions between different parts of the whole enhancing interpretation of data.

For Example:

Compare the sales of one region (quantitative value) to total sales of all regions (entire set of values).

Ratios are commonly used to measure changes over time.

  • Helps track and quantify fluctuations or trends in values over specific periods.

  • Useful for analysing trends and making comparisions in financial or performance data.

For Example:

Express the degree of change, such as comparing expenses from one month to the next.

Correlation

Correlation examines paired sets of quantitative values to determine if changes in one correspond to changes in the other.

Crucial as they enable predictions about how values in one variable may change based on knowing or influencing values in another variable.

For Example:

Consider whether there’s a relationship between employees’ job tenure and their productivity - does productivity increase, decrease, or show no significant change with tenure?

[Numbers that Summarise]

Statistics provide methods for summarising or aggregating data to effectively communicate quantitative messages.

Summarisation, also known as aggregation, involves condensing large datasets into key numerical representations.

Summarisation

Data summarisation includes:

Visual methods

Statistical methods

Focusing on lesser-known techniques beyond basic sums and counts

Example:

When reporting sales to executives, summarising data such as:

  • Total sales in rupees

  • Number of orders

  • Average order size

It is more effective than detailing each individual transaction.

Let’s Explore

Measures of Average:

  • Mean

  • Median

Measures of Variation:

  • Spread

  • Standard Variations

Measures of Correlation:

  • Linear Correlation Coefficient

Measures of Ratio:

  • Rate or percentage

[Measures of Average]

Average

An average is a single number that represents the center of an entire set of numbers.

Four distinct ways in statistics to measure the center of a set of numbers, and all of them are called averages:

  • Mean

  • Median

  • Mode

  • Midrange

Average

Useful to understand how these four differ.

Selecting the wrong type of average for your message could mislead your audience.

Mean

The term “average” is often used to refer to the arithmetic mean.

In software, the function calculating the mean is often labelled as “average” or “AVG”.

Statisticians prefer precise terminology, using “mean” instead of “average”.

The mean is calculated by summing all values and dividing by the number of values.

While simple to calculate, the mean isn’t always the best choice for representing data.

Mean - When its Not Good

The mean considers every value, including extremes, which can sometimes be misleading.

An example shows that the mean may give a misleading impression of typical salaries in a marketing department if extreme values are present.

  • The mean is much higher than most salaries, suggesting employees are better compensated than they actually are.

For typical values, a measure of centre that isn’t sensitive to extremes is more appropriate.

Median

The median is useful for telling quantitative stories because it is not sensitive to extremes.

Calculation of the median:

  • Sort the values in order.

  • Find the middle value.

For an even number of values

  • The median is the average of the two middle values.

  • Software or calculators can automatically handle median calculations.

Eg:

In a set of 13 sorted salaries, the median is 7th value ($62,000).

The median is the 50th percentile, meaning 50% of values fall below it.

The median better represents typical values than the mean in cases with extreme values.

The median is preferable for showing typical values, but the mean is better for summarising overall financial impact.

Mode

The mode is the value that appears most frequently in a data set.

In the above marketing department salaries example, the mode is $165,000 because it appears more than once.

Data sets can be:

  • No mode (if no value repeats)

  • Bimodal (if two values appear twice)

  • Multimodal (if more than two values appear with the same high frequency)

Modes are rarely useful for most data presentations.

Midrange

The midrange is the value midway between the highest and lowest values in a set.

Calculation:

  • Add the highest and lowest values and divide by two.

Quick to calculate but sensitive to extremes.

Less reliable than the mean or median unless values are uniformly distributed across the range.

[Measures of Correlation and Ratios]

Co-relate

A particular type of quantitative relationship where two paired sets of quantitative values are compared to one another to see if they correspond (i.e. co-relate) in some manner.

For example:

Does tenure on the job (number of years of working in an organisation) relate to productivity (say performance rating)?

Linear correlation coefficient answers:

Does a correlation exist?

If so, is it strong or weak?

If so, is it positive or negative?

Linear Correlation Coefficient

Measures the direction (positive or negative) and degree (strong or weak) of the linear relationship between two paired sets of quantitative values.

Paired sets of Quantitative Values:

Two sets of measurable values are considered when examining the relationship between two variables e.g. employee’s tenure and productivity.

Limitation:

Can not describe an inconsistent relationship.

Eg: productivity increases with tenure up to a point and then decreases.

Direction of Correlation

Direction of Correlation:

Can be positive or negative.

Positive Correlation:

As the value of one variable (A) increases, the value of the other variable (B) also increases, and vice versa.

Negative Correlation:

As the value of one variable (A) increases, the value of the other variable (B) decreases, and vice versa.

Correlation: Calculation and Interpretation

Calculation:

Manual calculation of the linear correlation coefficient involves several steps, but software or calculators usually perform this task.

Interpretation Focus:

Understanding the resulting value is more important than calculating it manually.

Correlation: Interpretation Guidelines

All values range between +1 and -1.

A value of 0 indicates no linear correlation.

A value of +1 indicates a perfect positive linear correlation.

A value of -1 indicates a perfect negative linear correlation.

The greater the value (positive or negative), the stronger the linear correlation.

Correlation: Interpretation

Measures of Ratio

Correlations (which measure relationships between multiple paired sets of values)

Ratios measure relationships between a single pair of values.

  • Rates and percentages are most concise and useful for tables and graphs.

For example:

  • Book-to-bill rate compares the value of booked sales orders to the value of actual billings generated.

Expression of Ratios

Four conventional ways:

As a sentence: eg “Two out of every five customers who access our website place an order.”

As a fraction (eg 2/5)

As a rate (eg 0.4)

As a percentage (eg 40%)

Examples of Conventional Ratios

Book-to-Bill Rate:

Typically expressed as a rate eg, 1.25, indicating that for every five orders booked, only four have been billed.

Profit Margin:

Typically expressed as a percentage eg, 25%, indicating that for every $100 of revenue, $75 goes toward expenses, leaving a profit of $25.

Measures of Money

When comparisions of monetary value are expressed across time, adjusting the value to account for inflation produces the most accurate results.

When reporting monetary values that combine multiple currencies, you must first convert them all into a common currency.

[Introduction to Jupyter Notebooks]

You can use Anaconda and open Jupyter Notebook.

Consider: Creating a folder.

Consider: Creating a notebook.

Note that Shift + Enter runs code in a cell.

Note that you can create Markdown cells as well, to add explanations. You can add pictures, equations, etc.

(This is a nice feature of Jupyter Notebooks).

Consider: Downloading a notebook (and sharing it).

Consider: Getting rid of a cell (using the cut icon).

Consider: Restart and clear output option.

Note that Jupyter Notebooks have extension .ipynb.

Consider: Google Colab, Microsoft Azure Notebooks

[Introducing IDEs and Code Editors]

IDLE and Python shell are suitable for small tasks but become problematic for larger projects.

An IDE or a good dedicated code editor can make coding more enjoyable.

Choosing the best tool depends on individual preferences and processes.

Will cover tools built exclusively for Python development and general development tools adaptive for Python.

IDEs (Integrated Development Environment)

Programs dedicated to software development.

Integrate several tools designed for software development.

Common tools include:

  • Code editors with syntax highlighting and auto-completion.

  • Build, execution, and debugging tools, Source control tools.

Support multiple programming languages and offer many features.

Typically large and take time to download and install.

May require advanced knowledge for proper use.

Dedicated Code Editors

Can be as simple as text editors with syntax highlighting and code formatting.

Good code editors can execute code and control a debugger.

The best code editors interact with source control systems.

Smaller and quicker compared to IDEs.

Often less feature-rich than IDEs.

Core Features for a Good Python Coding Environment

Save and Reload Code Files:

Essential for saving work and reopening it in the same state.

Run Code from Within the Environment:

Avoids the need to exit the editor to run Python code.

Debugging support:

Allows stepping through code as it runs, crucial for troubleshooting.

Syntax Highlighting:

Helps quickly identify keywords, variables, and symbols, making code easier to read and understand.

Automatic Code Formatting:

Recognises syntax rules (eg indenting after colons in loops) and formats code automatically.

Additional Desirable Features:

  • Source code control

  • Extension model

  • Build and test tools

  • Language help

Categories

General Editors and IDEs with Python Support:

IDEs:

  • Eclipse + PyDev

  • Visual Studio

Code Editors:

  • Visual Studio Code

  • Sublime Text

  • GNU Emacs

  • Vi / Vim

Categories

Python-Specific IDEs:

  • PyCharm

  • Spyder

  • Thonny

Visual Studio Code (Code Editor)

A full-featured code editor available for Linux, macOS, and Windows.

Small and lightweight yet powerful and extensible.

Built on Electron, which uses JavaScript, HTML, and CSS for desktop applications.

Open source and highly configurable for various tasks.

Pros:

Easy installation of Python support through the Marketplace.

Automatically recognizes Python installation and libraries.

Cross-platform availability (Linux, macOS, Windows)

Full-featured with a small footprint.

Open source and extensible.

Cons:

Not a native app due to being built on Electron.

Some users may avoid it due to its association with Microsoft.

Sublime Text (Code Editor)

Created by a Google engineer aiming to build a better text editor.

Extremely popular and supported on all platforms.

Built-in support for Python code editing.

Offers a rich set of extensions (called packages) to enhance syntax and editing features.

Pros:

Popular with a strong community following.

Fast, small, and well-supported as a code editor.

Rich set of extensions for customisation.

Cons:

Installing additional Python packages can be complex.

All packages are written inn Python, often requiring manual script execution within Sublime Text.

Not free, although the evaluation version can be used indefinitely.

No direct support for executing or debugging code within the editor.

PyCharm (IDE)

A dedicated IDE specifically for Python development.

Available in both paid (Professional) and free open-source (Community) editions.

Installs quickly and easily on Windows, macOS, and Linux.

Pros:

Supports Python development directly out of the box.

Allows running and debugging Python code directly within the IDE.

Includes support for source control and project management.

Considered the de facto Python IDE with extensive support and a strong community.

Cons:

Can be slow to load.

Default settings might require adjustments for existing projects.

Spyder (IDE)

An open-source Python IDE optimised for data science workflows.

Included with the Anaconda package manager distribution.

Targeted at data scientists using Python.

Pros:

Integrates well with common Python data science libraries like SciPy, NumPy, and Matplotlib.

Features a code editor with robust syntax highlighting. Python code completion, and an integrated documentation browser.

Includes a unique variable explorer for displaying data in a table-based layout within the IDE.

Offers IPython/Jupyter integration.

Free and open source.

Supports Windows, macOS, and Linux.

Cons:

May feel too basic for experienced Python developers for daily use.

Considered more of a special-purpose tool rather than a primary editing environment.

[Familiarising with Anaconda, Miniconda, and Conda]

Package, Environment and Package Manager

A package is a piece of code someone else has written which can be run and often serves a specific purpose. Can consider a package as a tool you can use for your own projects.

Environment is a collection of tools or packages which are managed separately for each project.

Package manager helps to take care of different packages by handling installing updating and removing them.

Why are Anaconda, Miniconda, and Conda important?

Experimental nature of Machine Learning and Data Science.

  • It involves a lot of trial and error, necessitating tools that support this iterative process.

Reproducibility:

  • When a successful experiment is found, tools help ensure that it can be reproduced reliably.

Collaboration:

  • Facilitate sharing work with colleagues, teams, or the wider community by providing a consistent environment.

What is Anaconda?

Anaconda is a Python distribution aimed at Data Scientists that includes 250+ packages.

  • With easy access to 7500+ additional packages

  • Value proposition is that one can download it (for free) and “everything just works”

  • Available for Mac, Windows, Linux

A new Anaconda distribution is released a few times a year.

  • Within each distribution, the versions of the included packages have all been tested to work together.

What is Conda?

An open source package and environment manager that comes with Anaconda and Miniconda.

As a package manager, can use Conda to install, update, and remove packages and their “dependencies”.

As an environment manager, can use Conda to manage virtual environments an alternative to virtualenv, pipenv, and other related tools.

What is Miniconda?

A Python distribution that only includes Python, conda, their dependencies, and a few useful packages.

A free minimal install for Conda.

A great choice if

  • prefer to only install the packages needed,

  • sufficiently familiar with Conda.

If need more packages, can use Conda to install from thousands of packages available by default in Anaconda’s public repo, or from other channels, like Conda Forge or Bioconda.

[Introduction to Python]

Program and Programming Language

Program: Set of instructions that perform a specific task executed by computer.

  • Required by computer to function.

Written by programmer using programming languages.

  • Like C, C++, Java, Python, etc.

Executed with compiler and interpreter.

Python’s Growth as a Programming Language

Python as a Programming Language

Designed by Guido van Rossum around 1990.

O-O rapid prototyping language.

Not just a scripting language.

Easy to learn, read, use.

Extensible (add new modules).

Highly readable.

Touchy Feel Properties

Open Source:

  • Copyrighted but use not restricted.

  • Owned by independent non-profit, PSF.

Mature (34 years old).

Supportive user community

  • Plenty of good books.

  • Too Active user community.

Simple design, easy to learn:

  • Reads like “pseudo-code”.

  • Suitable as first language.

  • Suitable as last language.

Python Applications

Use Python for:

Web Development: Django, Pyramid, Bottle, Tornado, Flask, web2py

GUI Development: tkInter, PyGObject, PyQt, PySide, Kivy, wxPython

Scientific and Numeric: SciPy, Pandas, iPython

Software Development: Buildbot, Trac, Roundup

System Administration: Ansible, Salt, OpenStack

Components of Python World

Core Python

Distributions

Frameworks / IDEs

Third Party Libraries

Python Ecosystem

Core Python:

Programming Language itself.

Some standard modules are available.

Other packages needs to be explicitly installed.

Python Distribution:

Python + Packages.

Majority of packages, libraries are already available.

Package management is simplified.

Anaconda from Continuum Analytics.

iPython and its iPyKit variant.

Frameworks / IDEs:

Use frameworks to create code and develop applications.

Provides a defined structure to the developers so that they can focus on the core logic of the application rather than on other elements.

Python web framework:

django, web2py, flask

Python IDEs:

IDLE, PyCharm, Spyder, Jupyter

Python Installation

Install Python directly

Use Python distribution

Use cloud based services

Install Python directly:

Install the Python language with installer.

Need to install other packages explicitly using pip install.

https://www.python.org/downloads/

Use Python distribution:

The open-source Anaconda.

Distribution is the easiest way to perform Python coding.

Works on Linux, Windows, and Mac OS X.

https://www.anaconda.com/distribution/

Use Cloud based services:

The simplest of all but needs Internet connectivity to use.

Code works in cloud service provided data center.

Microsoft Azure Notebooks

https://notebooks.azure.com/

Google Colaboratory

https://colab.research.google.com

[I/O Statements]

Python - Output Statements

print("Hello World from Python!")

Output: Hello World from Python!

print("Multiline output \n can also be printed.")

Output:

Multiline output 
 can also be printed.

print('String concatenation', 'also', 'works.')

Output:

String concatenation also works.

Python Outputs

print('3+4')

Output: 3+4

print(3+4)

Output: 7

print('3 + 4 = ', (3+4))

Output: 3 + 4 = 7

With a default separator:

print("This", "is", "first", "line")

Output: This is first line

Optional separator:

print("This", "is", "first", "line", sep = ", ")

Output: This, is, first, line

We can also use end

print("this is first line", end=".")
print("this is second line") 

Output: this is first line.this is second line

Python - Input Statements

input('What is your name?)

Output: What is your name?[Prompt box]

The user input can be saved as part of a variable.

name = input('What is your name?')
print("Welcome", name)

Output (if I enter Karthik):

What is your name?Karthik
Welcome Karthik

type(name)

Output: str

We can use type casting functions.

age = input("What is your age?")
print("You have entered age as ", age)
print("Type : ", type(age)) 

Output:

What is your age?25
You have entered age as  25 
Type :  <class 'str'> 
age = input("What is your age?")
modified_age = int(age) 
print("You have entered age as ", modified_age)
print("Type : ", type(modified_age))

Output:

What is your age?25 
You have entered age as  25
Type :  <class 'int'>
salary = input("What is your salary?")
modified_salary = float(salary) 
print("You have entered salary as ", modified_salary)
print("Type : ", type(modified_salary)) 

Output:

What is your salary?10.5
You have entered salary as  10.5 
Type :  10.5

The eval function converts the text entered by user into a number (int, float as required I guess) without explicitly typecasting.

[Basic Constructs]

Python - Modules and Packages

A module is a piece of software that has a specific functionality. For example, math module has functionality related to mathematical operations. In Python, A module is a file containing Python definitions and statements. The file name is the module name with the suffix .py appended.

As your program gets longer, you may want to split it into several files for easier maintenance. You may also want to use a handy function that you’ve written in several programs without copying its definition into each program.

To support this, Python has a way to put definitions in a file and use them in a script or in an interactive instance of the interpreter. Such a file is called a module; definitions from a module can be imported into other modules or into the main module (the collection of variables that you have access to in a script executed at the top level and in calculator mode).

Built in Python Modules

Python comes with a library of standard modules. In order to make use of the modules, they need to be imported into the script or notebook first. Then the definitions from that module can be used within the program.

import sys 
print(sys.version) 

Output (in the video):

3.11.5 | packaged by Anaconda, Inc. | (main, Sep 11 2023, 13:26:23) [MSC v.1916 64 bit (AMD64)]
import math 
math.sqrt(5) # compute the square root of a number 

Output:

2.23606797749979
import time 
time.sleep(5) # halts the execution for 5 seconds 

(After 5 seconds the cell execution is complete)

Ways of importing modules

Here are several ways to import some functions from Random module.

from random import randint 
from random import *
import random

1 First way imports just function randint from the random module.

2 Second way imports all the functions from the random module. Generally to be avoided as some names in module might interfere with your own variables.

3 Third way import module in a way that will not interfere with your variables names.

Using Alias

The as keyword can be used to refer the module with different name. Then use that alias to call the functions from that module.

import numpy as np 
np.arange(5) #creates array of 5 elements from 0 to 4

Output:

array([0, 1, 2, 3, 4])

Getting help

To get help on module, import it and then use dir(module_name) to get the list of functions defined in that module.

dir(np) 

Output: `A list of functions

help(np.arange)

Output: Details about np.arange method

Packages

Packages are namespaces which contain multiple packages and modules themselves. They are simply directories, but with a twist.

Each package in Python is a directory which MUST contain a special file called init.py. This file can be empty, and it indicates that the directory it contains is a Python package, so it can be imported the same way a module can be imported.

Packages are a way of structuring Python’s module namespace by using ‘dotted module names’. For example, the module name A.B designates a submodule named B in a package named A. Just like the use of modules saves the authors of different modules from having to worry about each other’s global variable names, the use of dotted module names saves the authors of multi-module packages like NumPy or Pillow from having to worry about each other’s module names.

Python - Data Types and Type Casting

Data Types

There are many different types of objects in Python. The most common object types: strings, integers, and floats.

Way to define literals

1 # integer 
1.5 # float
'my string'
True

Way to check type of literal

type(1) 

Output: int

type(1.0) 

Output: float

type(True)

Output: bool

Type casting

int(1.5) 

Output: 1

type(int(1.5)) 

Output: int

int('1') 

Output: 1

Python - Variables, Expressions & Statements

Variables

The major purpose of variable is to store a value from one part of program so that it can be used in other part of program. It can be used to store the value of computation that is carried out duing the flow of the program. A variable is a name that refers to a value.

Assignment statement

Assignment statement creates a new variable and store a value in it.

Eg:

message = "this is first string variable"

# a string type variable named message is created and value is stored in it 

print(message)

Output:

this is first string variable
num = 5 

# an int type variable named num is created and 5 is stored as value in int

print("num contains ----> ", num) 

print("type of num -----> ", type(num))

Output:

num contains ----> 5 
type of num -----> <class 'int'>
float_num = 3.5 

# a variable of type float is created and value 3.5 stored in it 

print("float_num contains ----->  ", float_num)

float_num = 4.5 

# value of variable can be changed 

print("float_num now contains ----->  ", float_num)

Output:

float_num contains ----->   3.5 
float_num now contains ----->   4.5

Variable naming

Choose the variable name that is meaningful and improves the readability of the program.

Rules to be kept in mind while naming the variables:

  • Letters, numbers, and underscore are allowed

  • Can not contain spaces

  • Can not start with number

  • It’s case sensitive

  • It should not be from the reserved keyword list

  • Variable name cannot be a keyword (like class)

Statements

It is unit of code that Python interpreter executes. It can be a single statement or group of statement like statements within loops, functions, etc. When interpreter hits statement, it executes the statement and displays the result (if there is any).

# function call is statement 

print('this is statement') 

Output: this is statement

# Assignment is statement but does not produce any output 

x = 5
# Compound statement - for loop 

for i in range(5): 
    print("Coming inside the loop ----> ", i + 1, " times") 
    print("i is ---> ", i)

Output (I wrote dots instead of the mid part):

Coming inside the loop --->  1  times 
i is --->  0 
. 
.
.
Coming inside the loop --->  5  times 
i is --->  4

Expressions

Expressions in Python can include operations among compatible types (eg integers and floats). For example, basic arithmetic operations like adding multiple numbers.

35 + 45 + 23

Output: 103

+ addition 
- subtraction 
* multiplication 
/ regular division 
// integer division # x = (x // y) * y + (x % y) 
% modulo
** exponent

[Data Structures - Strings and Tuples]

Python - Strings

Strings are sequence of characters. It can hold zero or more than zero characters in it. String is an object in Python. Python provides many functions to work on with the strings.

Creation

"this is string" # creates a string 
my_str3 = """string is spread 
across multiple lines""" # triple quotes can be used to create multi line strings 
type(my_str3) 

Output: str

len('this is string')

Output: 14

Indices

String characters has index associated with it starting from zero to len(string) - 1. Individual characters can be accessed by indices. Negative indexing is also available.

my_string = 'this is my string'
my_string[0] # accessing first character of string 

Output: t

my_string[1] # accessing second character of string 

Output: h

length = len(my_string) 
my_string[length - 1] 
# accessing last character of string 

Output: g

my_string[-1] # accessing last character of string using negative indexing 

Output: g

my_string[-2] 

Output: n

Slicing

Slice is portion of string accessed using the indices. [] operator allows us to extract part of string using the indices.

my_string = 'this is my string'
my_string[ : ] # extracts complete string 

Output: 'this is my string'

my_string[ 0 : 3 ] # extracts first three characters of string 

Output: 'thi'

The in operator

The in operator is helpful to find out if a string contains some other character or string.

my_string = "this is string"
'i' in my_string # check whether i is present in string

Output: True

'i' not in my_string # check whether i is not present in string 
'this' in my_string

Immutability

my_string = 'this is my string'
my_string[0] = 'p' 

# not allowed, as strings are immutable, i.e. cannot be modified

Output: TypeError: 'str' object does not support item assignment

my_string = 'this is new string'

# you can assign some different value to the string variable but its currently assigned value can not be changed 

Btw one can concatenate strings using +.

Looping

One can iterate over the characters present in the string using the for loop.

my_string = "python is beautiful" 
for str_char in my_string:
    print(str_char)

Output:

p
y
t
h
o
n

i
s

b
e
a
u
t
i
f
u
l

Python - String Operations

String object supports lot of methods. Lets explore some of them here.

dir(str) # list the methods of string 

String methods does not change the original string. If the change needs to be captured, it has to be assigned back to some variable.

my_string = 'Python is beautiful!'
my_string.lower() 

#converts to lowercase 

Output: 'python is beautiful'

my_string # original string is not altered

Output: 'Python is beautiful'

String methods

lower() - converts every character of string in lower case

upper() - converts every character of string in upper case

islower() - determines whether each character is in lowercase or not

Eg: my_string.islower()

isupper() - determines whether each character is in uppercase or not

Eg: my_string.isupper()

count() - counts the number of occurences of characters in string

Eg: my_string.count('t')

index() - returns the index of a character

Eg: my_string.index('t')

isalpha() - determines whether a character of string is letter or not

Eg: my_string[0].isalpha()

isdigit() - determines whether a character of string is digit or not

Eg: my_string[0].isdigit()

String stripping

Sometimes the strings come with whitespaces attached at both ends. The characters from the left and right side of string can be removed with the strip function.

stmt = "Chennai Super Kings are going to win IPL this time.    "

stmt.rstrip()

Output:

'Chennai Super Kings are going to win IPL this time.'
stmt = "    Chennai Super Kings are going to win IPL this time."

stmt.lstrip()

Output:

'Chennai Super Kings are going to win IPL this time.'
stmt = "    Chennai Super Kings are going to win IPL this time.    "

stmt.strip() 

Output:

'Chennai Super Kings are going to win IPL this time.'

Substrings

The strings which are part of string are substrings. For eg, ‘beautiful’ is substring of string ‘python is beautiful’. There are several functions to deal with substrings.

my_string = 'python is beautiful'
my_string.find('is') 

Output: 7

my_string.replace("t", "s") 

Output: 'pyshon is beausiful'

We also have startswith() and endswith().

String Splitting

my_string = "Python is beautiful" 
my_string.split() 

# split using default delimiter i.e. whitespace 

Output: ['Python', 'is', 'beautiful']

my_string.split(' is ') 

# split using defined delimiter i.e. 'is'

Output: ['Python', 'beautiful']

Python - Tuples

Tuple is data structure that holds collection of different objects in it. For example, a tuple can hold numerical as well as string values in it along with the other collections like list, tuples, etc.

Creation

Empty tuple can be created with angular brackets.

Tuple elements are enclosed within angular brackets.

my_tuple = ('a', 'b', 'c', 'd', 'e') 
my_tuple

Tuple can hold elements belonging to different data types.

my_mixed_tuple = ('a', 1, 1.1, True)
my_mixed_tuple

Type of tuple variable is tuple.

print(type(my_tuple)) 

Output: <class 'tuple'>

Size of the tuple can be determined using the len() function.

len(my_tuple)

Output: 5

Indexing

Tuple elements can be accessed using the bracket operator [].

Slicing

Like with strings. For eg, my_mixed_tuple[0 : 2].

Concatenation

Two tuples can be joined using the + operator.

Immutability

Tuple elements can not be altered (we already saw this with strings).

my_tuple = ('a', 'b', 'c', 'd', 'e') 
my_tuple[0] = 'p' 

# throws error as individual tuple elements can not be altered 

New elements can not be added to the tuple using append()

Operations

Index, Count.

Eg:

my_tuple.index('c')
my_tuple.count('a') 

Iterable

Tuple is iterable.

for i in range(len(tuple)): 
    print("(", i, ") ", my_tuple[i])

Sorting

scores = (10, 8, 3, 2, 4, 0, 0, 3) 
scores 

Output: (10, 8, 3, 2, 4, 0, 0, 3)

sorted(scores)

Output: [0, 0, 2, 3, 3, 4, 8, 10]

(The output seems to be a list)

scores 

Output: (10, 8, 3, 2, 4, 0, 0, 3)

[Data Structures - Lists]

Python - Lists

Lists are ordered collection of elements. The elements can be of any type like strings, numbers, booleans, lists, tuples, etc. Each element of list has an address associated with it i.e. index. This index can be used to get access to the list elements.

Creation

Empty list can be created using [] (Usually [] brackets are used to define a list)

List can have multiple elements in it.

my_number_list = [1, 2, 3, 4, 5, 6, 7]
my_number_list 

But it can hold elements of different kinds as well.

my_mixed_list = [1, 1.1, 'a', True]
my_mixed_list

Type can be checked with type() function.

type(my_mixed_list)

Output: list

Length of list can be determined using len() function.

Eg: len(my_number_list)

List can be printed using the print() function.

Eg: print(my_number_list). Output: [1, 2, 3, 4, 5, 6, 7]

Concatenation

+ can be used to concatenate lists.

Indices

Index is the location where the list element is present in the list. List elements are ordered. Indexing starts from 0.

Eg: my_list[0] # access first element of the list

Slicing

Slice is portion of list accessed using the indices. [] operator allows us to extract part of list using indices.

Eg: my_list[ 0 : 3 ]. Output: [1, 1.1, 'a']

The in operator

The in operator is helpful to find out if a list contains an element in it or not.

Eg: 1 in my_list # check whether 1 is present in list

Looping

One can iterate over elements present in the list using for loop.

Eg:

for element in my_list: 
    print(element) 

Built-in functions

sum() can be used to get sum of all elements in list.

min() can be used to get minimum value out of all elements in list.

max() can be used to get maximum value out of all elements in list.

List aliasing

The same list can be referred by two different list variables. Changing one list affects the other as well.

list1 = [1, 2, 3, 4, 5]

list2 = list1 

# Alter the elements of list1, both lists will get altered

# Alter the elements of list2, both lists will get altered 

List copying

In order to avoid this side effect, list elements can be copied using the following way. By that way, alterations in one list does not affect the elements in the other list from which its copied.

There are three ways of list copying:

new_list = orig_list.copy()

new_list = orig_list[:]

new_list = list(orig_list) 

Python - List Operations

Operations on lists

Lists are mutable, i.e. the elements can be added, altered or removed from the list.

my_list = [1, 2, 3, 4, 5]
my_list 

append() can be used to add an element to the list.

my_list.append(6) # add new element 6 at end of list 
my_list

Output: [1, 2, 3, 4, 5, 6]

my_list.append([7, 8]) # add list as new element in list 
my_list 

Output: [1, 2, 3, 4, 5, 6, [7, 8]]

extend() can also be used to add list of elements into list.

my_new_list = [1, 2, 3, 4, 5] 
my_new_list
my_new_list.extend([6]) 
my_new_list

Output: [1, 2, 3, 4, 5, 6]

my_new_list.extend([7, 8, 9])
my_new_list

Output: [1, 2, 3, 4, 5, 6, 7, 8, 9]

insert() can be used to add an element at the desired position.

my_list = [1, 2, 3, 4, 5, 6] 
my_list
my_list.insert(0, 'a') 

# insert 'a' at first position 

my_list

Output: ['a', 1, 2, 3, 4, 5, 6]

my_list.insert(len(my_list), '*****') 

# insert '*****' at last position 

my_list 

Output: ['a', 1, 2, 3, 4, 5, 6, '*****']

remove() can be used to find the first occurence of the element and to remove it.

my_list = [1, 2, 3, 4, 5] 
my_list 
my_list.remove(3) 

# find first occurence of '3' and remove it from the list 

my_list 

Output: [1, 2, 4, 5]

pop() can be used to remove an element from specified index and store its value in a variable.

my_list = [1, 2, 3, 4, 5] 
my_list 
popped_element = my_list.pop(4) 
print("list--->", my_list) 
print("popped element--->", popped_element) 

Output:

list---> [1, 2, 3, 4]
popped element---> 5

del() can also be used to remove element from the list at specified location.

my_list = [1, 2, 3, 4, 5]
my_list
del(my_list[1]) 

# removes the element at 1 index 

my_list 

Output: [1, 3, 4, 5]

Assignment operator can be used to alter the element of list at the desired index.

my_list = [1, 2, 3, 4, 5] 
my_list 
my_list[0] = 1111 

# change the element at first position to 1111

my_list 

Output: [1111, 2, 3, 4, 5]

List Methods

sort() can be used to sort the list elements. Default is ascending order.

my_list = [11, 2, 39, 4, 57] 
print("list before sorting--->", my_list)

print()

my_list.sort()
print("list after sorting--->", my_list)

Output:

list before sorting---> [11, 2, 39, 4, 57]

list after sorting---> [2, 4, 11, 39, 57]

Sort order can be specified as argument. Default is reverse = False. Arranges elements in descending order.

my_list = [11, 2, 39, 4, 57]
print("list before sorting--->", my_list)

print()

my_list.sort(reverse = True)
print("list after sorting--->", my_list)

Output:

list before sorting---> [11, 2, 39, 4, 57]

list after sorting---> [57, 39, 11, 4, 2]

Alternate way to get the sorted list using sorted() method. Original list is not modified.

my_list = [11, 2, 39, 4, 57]
print("list before sorted--->", my_list) 

print()

print("list after reversing the order--->", list(sorted(my_list)))

print()

print("original list after using sorted function--->", my_list)

Output:

list before sorted---> [11, 2, 39, 4, 57]

list after reversing the order---> [2, 4, 11, 39, 57]

original list after using sorted function---> [11, 2, 39, 4, 57]

reverse() can be used to reverse the order of list elements.

my_list = [11, 2, 39, 4, 57] 
print("list before reversing the order--->", my_list) 

print() 

my_list.reverse() 
print("list after reversing the order--->", my_list) 

Output:

list before reversing the order---> [11, 2, 39, 4, 57]

list after reversing the order---> [57, 4, 39, 2, 11]

Alternate way to get the reversed list is using reversed() method. Original list is not modified.

my_list = [11, 2, 39, 4, 57]
print("list before reversing the order--->", my_list) 

print() 

print("list after reversing the order--->", list(reversed(my_list)))

print() 

print("original list after using reversed function--->", my_list)

Output:

list before reversing the order---> [11, 2, 39, 4, 57]

list after reversing the order---> [57, 4, 39, 2, 11]

original list after using reversed function---> [11, 2, 39, 4, 57]

index(x) can be used to determine the location of first occurence of element x in the list.

my_list = [11, 2, 39, 4, 57] 
my_list.index(39) 

# find the location at which element 39 first occurs 

Output: 2

count(x) can be used to determine the count of occurences of element x in the list.

my_list = [11, 2, 39, 4, 57, 39, 34, 39]
my_list.count(39) 

# count the number of times 39 appears in the list 

Output: 3

comments powered by Disqus