# DataScientYst - Data Science Simplified > DataScientYst - Data Science Tutorials, Exercises, Guides, Videos with Python and Pandas Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages ### Privacy Policy URL: https://datascientyst.com/privacy/ Last updated: 2022-08-30T20:24:12.000Z This Privacy Policy explains how information is collected, used and disclosed by DataScientyst with respect to user’s access and use of our service through the application (Referred to below as “DataScientyst”). ### 1\. Information collection When using DataScientyst, we ask certain information from you: **“Personal Information”** We dont collect any personal data, except that we use 3rd party statics for stat purpose. Users who contact us via email, the email addresses and information you submitted voluntarily will also be collected. ### 2\. Information usage Your information will not be shared with others and is only used internally for the purposes described below: - to provide our services or information you request, and to process and complete any transactions; - to respond to your emails, submissions, questions, comments, requests, and complaints and provide customer service; - to analyze usage and trends with anonymous user data, and to improve the quality of our service and user experience; - to send you confirmations, updates, security alerts, and support and administrative messages and otherwise facilitate your use of, and our administration and operation of, our services; Your location data will not be shared with others and is only used internally for the purposes described below: - to provide you reminders based on locations; ### 3\. Cookies Cookies are required on this Website. We use them to collect visitors preferences and thus to better optimize the user experience. Users can disable cookies in their own browser settings, but please note that you may not be able to access certain features on our Website as a result. ### 4\. Changes This Privacy Policy may be revised and modified at some time in the future. We will post on this page and notify you via official notifications or emails. Please check back periodically to keep informed of updates or changes to this Privacy Policy. By continuing to access and to use DataScientyst, you are agreeing to be bound by the revised policy. ### About us. URL: https://datascientyst.com/about/ Last updated: 2022-08-30T20:24:39.000Z I'm Johny D K, a software engineer, data scientist and dreamer. Hey there! I'm a data scientist from East Europe. This is my place on the web for data science projects, tutorials and guides. I love facing data challenges and finding hidden insights and relations. I've started as a database administrator long time ago. My first database was Oracle. I remember the hard times which I had with SQL. This period was for one year and something. Next challenge was ETL, data mining and data migration. At that time I was sleeping with financial data literally. Still having problems with data manipulation, SQL, excel tables... but I realized that the right career for me is Data Scientist. ![panda_about](https://datascientyst.com/content/images/2022/06/panda_about.png) Luckily for me the hype for Data Science grow and I was able to move to the next level - Python, Pandas, R.. This was huge step ahead. Another big change at that time was Linux. Initially I was doing a task for weeks, later the weeks become days and finally hours. And even now I'm fascinated how much I can learn every single day :) ## Hobbies The list of my hobbies is long so I'll keep to the most important ones. The first and biggest one is nature. I like everything related to nature - walking, learning, breathing. Long walk in the forest is the most inspirational thing for me. I like sports like swimming, running, fitness. They help me to relax after hard data challenges. The active days are the only rest for my brain. I love reading science fiction, space, biotech, geography. This is my reward at the end of the day. ## Side Projects One of my first projects is [fantasyan](https://fantasyan.com/?ref=datascientyst.com). Combination of my fantasies, AI writer and analysis of science fiction books. Recently I don't have much time to advance with this project. **Genome Project** is another interesting project for me. I'm helping biologist in finding patterns in human DNA. Oh yeah, it's not easy. [Softhints](https://softhints.com//?ref=datascientyst.com) is a place where I share my experience related to programming, Python, Linux, Automation and more. Thanks to my readers and viewers for the nice words! [Softhints YouTube Channel](https://www.youtube.com/channel/UCg5rvP%5FD735oSBatdcH5ZFA/?ref=datascientyst.com). ## Favourite Quote One of my favourite quotes is from book Dune: > “Muad'Dib learned rapidly because his first training was in how to learn. And the first lesson of all was the basic trust that he could learn. It's shocking to find how many people do not believe they can learn, and how many more believe learning to be difficult. Muad'Dib knew that every experience carries its lesson.” ## Books ### Data Science & Programming - [Python Data Science Handbook](https://jakevdp.github.io/PythonDataScienceHandbook//?ref=datascientyst.com) ### Fiction - [Dune](https://www.goodreads.com/book/show/44767458-dune?from%5Fsearch=true&from%5Fsrp=true&qid=Yt1BrfdRMm&rank=1&ref=datascientyst.com) - [The Sky Lords](https://www.goodreads.com/book/show/1825519.The%5FSky%5FLords/?ref=datascientyst.com) - [Musashi](https://www.goodreads.com/book/show/1825519.The%5FSky%5FLords?ref=datascientyst.com) ### Data Science Challenges URL: https://datascientyst.com/data-science-challenges/ Last updated: 2021-11-16T15:28:03.000Z Solving **Data Science challenges are a great way to practice and improve your analytical skills.** Data Science challenges are small projects that exercise your skills exploring, processing and analyzing data. In this article, we'll add Data Science challenges that you can solve to test and develop your skills in different scientific areas like: - programming - math - biology On a weekly or monthly basis we will add challenges. Feel free to contribute with ideas for challenges. ## Challenge 1: Data Styling Can you style DataFrame like a master? [Data Science Challenge 1: Data Styling ](https://datascientyst.com/data-science-challenge-data-styling/) ### Data Science Tutorial in Python and Pandas URL: https://datascientyst.com/data-science-tutorial-python-pandas/ Last updated: 2021-11-16T13:47:01.000Z Pandas is a mature, powerful, open source and highly flexible Python library focused on data analysis and manipulation. Pandas is one of the most popular tools for doing Data Science Some of the key benefits of Pandas are that: - it's easy to start and learn - open source - efficient for variety of datasets and file types - extensive set of features - high quality resources Another strong point is that, while Pandas is quite mature and well-established, it's very actively maintained and has a thriving dev community. This makes it quite up to date and aligned with the Python and Data Science ecosystem right now. Of course, there's a lot to learn to work confidently with Pandas. In the next months each month we will publish new chapter of Data Science Tutorial in Python and Pandas Let's get started... Soon :) ### Thank you very much! URL: https://datascientyst.com/thank-you/ Last updated: 2022-08-11T06:41:38.000Z All the content on the DataScientYst.com will be available for free for everyone. You can support our work by subscribing for the paid membership. This will allow us creating new content and fresh ideas related to data science. Thank You! ![](https://datascientyst.com/content/images/2022/08/thank-you-7377027_640.png) P.S. Your opinion is important for us. Feel free to reach about new ideas or content here: [Improve DataScientyst](https://docs.google.com/forms/d/e/1FAIpQLSc7rHn8Dh-Zu8twMDDFcYrLrK1kps2NlPuZcEHVcYAYWhXCMA/viewform?ref=datascientyst.com) ### Data For AI URL: https://datascientyst.com/data-for-ai/ Last updated: 2026-07-22T21:53:15.000Z ## 1\. What is Data for AI? High-quality data is the foundation of effective AI. From customer interactions and operational records to images, documents, and sensor data, AI systems rely on diverse and accurate information to learn, make predictions, and generate insights. The better the data, the more reliable and valuable AI outcomes become. ### 1.1 AI-Ready Data AI-ready data is clean, structured, well-governed, and accessible for machine learning and generative AI applications. It has been prepared through processes such as validation, labeling, standardization, and enrichment, ensuring that AI models can use it efficiently while maintaining quality, consistency, and compliance. ### 1.2 Trends Organizations are increasingly investing in data quality, governance, and real-time data pipelines to support AI initiatives. Key trends include the rise of synthetic data, automated data preparation, vector databases for generative AI, and stronger emphasis on privacy, security, and responsible AI practices as businesses scale AI across their operations. ### 1.3 Sources AI data comes from a wide range of sources, including enterprise databases, customer interactions, IoT devices, social media, public datasets, documents, images, videos, and third-party providers. Combining multiple data sources helps create more comprehensive and accurate AI models. ## 2\. Types AI systems use different types of data depending on the application. Structured data, such as tables and spreadsheets, is organized and easy to analyze, while unstructured data includes text, images, audio, and video. Semi-structured data, such as JSON and XML files, combines elements of both, providing flexibility for modern AI and analytics workflows. ### 2.1 Video & Audio Data Find and process large-scale video and audio datasets to train multimodal AI models. Rich multimedia data enables AI to better understand speech, visual content, context, and interactions across a wide range of applications. ### 2.2 Web Data Extract structured and unstructured data from websites to power AI model training, analytics, and retrieval-augmented generation (RAG) pipelines. Web data provides access to diverse, real-time information from across the internet. ### 2.3 Web Index Data Find pre-indexed, continuously updated web content for fast information retrieval in RAG systems and AI agents. Web index data reduces latency and provides scalable access to current web knowledge without requiring custom web crawling. ### What can I make for you, today? URL: https://datascientyst.com/service/ Last updated: 2026-07-22T21:47:52.000Z ## Does this sound familiar? - Your **data is scattered** across multiple systems, making it difficult to **trust your reports** and make **confident business decisions**. - You're planning a **data migration**, but you're worried about **data loss**, **downtime**, or **costly mistakes**. - **Poor data quality**—duplicates, missing values, and inconsistent records—is slowing your business down. - You want to leverage **Data Science** and **AI**, but your data isn't **clean, reliable, or analytics-ready**. - You need an experienced partner to **migrate your data**, **improve data quality**, and **transform your data into a strategic asset**. ## Meet Your Data Consultant Hi, I'm **John**—a **Data Science, Data Migration, and Data Quality Consultant** helping organizations unlock the full value of their data. With experience delivering **data migrations**, improving **data quality**, and building **analytics solutions**, I help businesses transform fragmented, unreliable data into trusted, actionable insights. I've helped organizations: - **Migrate** data securely between legacy and modern platforms. - **Improve data quality** by eliminating duplicates, inconsistencies, and missing data. - **Build** reliable data pipelines and reporting solutions. - **Develop** data science and analytics solutions for smarter decision-making. - **Create** scalable data strategies that support AI and business growth. Every organization has unique data challenges. I work closely with your team to identify opportunities, minimize risks, and deliver practical solutions that create lasting business value. Whether you need a **successful data migration**, **trusted data**, or an **AI-ready data strategy**, I'm here to help. ## Posts ### Positron: A New Data Science IDE Worth Knowing About URL: https://datascientyst.com/positron-a-new-data-science-ide-worth-knowing-about/ Last updated: 2026-08-29T13:13:31.000Z If you spend your days moving between Python and R, wrangling dataframes, or switching between RStudio and VS Code depending on the task, there's a new tool worth putting on your radar: **Positron**, a free IDE built specifically for data science. It's new to me, and relatively new itself — dating from around mid-2024\. I learned about it by checking out the GitHub profile of Wes McKinney, author of pandas. I'd been searching for a good data science IDE for a long time, so I decided to give it a chance. ![](https://datascientyst.com/content/images/2026/08/positron-data-science-ide.webp) ## What is it? Positron is a desktop IDE from **Posit** (formerly RStudio) designed around one idea: > data science work shouldn't force you to pick a single language or bounce between exploration tools and production tools. It's built as a fork of VS Code, so it inherits a familiar editor, extension ecosystem, and command palette, but it's been reshaped around the classic four-pane data science workflow that longtime RStudio users will recognize — console, editor, variables, and plots — while adding native support for both Python and R side by side. **Contributor spotlight: Wes McKinney.** One of the most notable people behind Positron is Wes McKinney, the creator of **pandas**, co-creator of **Apache Arrow**, and author of *Python for Data Analysis*. McKinney joined Posit as Principal Architect specifically to represent the needs of the Python data ecosystem inside a company that historically focused on R. He's been central to Positron's design, including its Data Explorer component for quickly inspecting CSV and Parquet files, and has described the project as essentially "a next-generation RStudio that works for R and Python." Positron is free and source-available under the Elastic License 2.0\. An enterprise version, Positron Pro, is also available through Posit Workbench for teams that need centralized management, SSO integration, and scalable compute. ## How to install it? Installation is straightforward on all major platforms: - **Windows**: Download the `.exe` installer (user or system install, x64 or ARM64) from the [Positron download page](https://positron.posit.co/download.html?ref=datascientyst.com). Make sure you have the latest Visual C++ Redistributable installed. - **macOS**: Download the `.dmg` for Apple Silicon or Intel. - **Linux**: Download the `.deb` (Ubuntu/Debian) or `.rpm` (Red Hat/Fedora) package for x64 or ARM64. Positron runs without any language installed, but if you plan to use Python or R: - **Python**: Positron works with any actively supported Python version. It supports `venv`, `uv`, `pyenv`, and `conda`/`pixi` environments — and can install a Python version for you automatically via `uv` if none is found. - **R**: You'll need R 4.2 or higher. If you manage multiple R versions, the [rig](https://github.com/r-lib/rig?ref=datascientyst.com) tool works well alongside Positron. Once installed, Positron checks for updates automatically. ## Installation issues with Ubuntu, Linux Mint Linux users — particularly on Ubuntu and Linux Mint — have occasionally run into dependency errors when installing the `.deb` package through a graphical installer like GDebi or the Mint Software Manager, typically a complaint about an unsatisfied `libglib2.0` dependency. This is documented in [Positron GitHub issue #6162](https://github.com/posit-dev/positron/issues/6162?ref=datascientyst.com). The practical fix reported by users experiencing this is to skip the GUI installer entirely and install the package directly from the terminal, letting `dpkg`/`apt` resolve dependencies on their own rather than relying on the GUI tool's dependency checker, which can be overly strict or outdated on Mint. In short: ```bash sudo dpkg -i Positron---x64.deb ``` If `dpkg` reports missing dependencies, follow up with: ```bash sudo apt --fix-broken install ``` Alternatively, you can point `apt` directly at the local file, which handles dependency resolution in one step: ```bash sudo apt install ./Positron---x64.deb ``` Several Mint users have confirmed this resolves the install cleanly — the issue turned out to be with the GUI `.deb` installation tools on Mint, not with the Positron package itself. ## What is supported Positron's feature set spans the full data science workflow: - **Languages**: First-class support for Python and R in a single editor, with the team designing it to be polyglot from the ground up (support for languages like Julia has been discussed as a future possibility). - **Data Explorer**: Fast, interactive inspection of large dataframes, CSVs, and Parquet files — click a file and it opens in a dedicated data grid with filtering and summary statistics. - **Variables and Plots panes**: For inspecting objects and visual output without leaving the IDE. - **Notebooks**: A built-in notebook editor that bridges Jupyter functionality with IDE features. - **Quarto**: Native support for building reproducible reports and presentations. - **Connections pane**: Explore and query database schemas and tables interactively. - **AI-assisted features**: Integrated AI assistance contextualized for data science tasks, including code completion and chat. - **Remote work**: Remote SSH support and integration with Posit Workbench for team-scale, server-based development. - **VS Code compatibility**: Because it's a VS Code fork, many existing extensions and keybindings carry over, and there are dedicated migration guides from both VS Code and RStudio. ## Who is it useful for? Positron is aimed squarely at: - **Data scientists and analysts** who work in Python, R, or both, and want one tool instead of switching between RStudio and a separate Python IDE. - **RStudio users** looking for a modern, actively developed successor with better performance on large datasets and native Python support. - **VS Code users** in data-heavy roles who want data science-specific panes (Variables, Data Explorer, Plots, Connections) without assembling their own extension stack. - **Teams and enterprises** that want centralized IDE management across Positron, RStudio, VS Code, and JupyterLab via Posit Workbench. It's less suited to general-purpose software engineering outside the data science context — for that, plain VS Code or another general IDE is still the more natural fit. ## Conclusion Positron feels like a natural response to how data science is evolving. As AI creates new demands for data exploration, analysis, and experimentation, data scientists increasingly need tools that can keep up without forcing them into a single language or workflow. Positron brings Python and R together in one environment, while supporting multiple operating systems and making it easier to move from exploration to production. It also benefits from a strong foundation. With Wes McKinney — the creator of pandas and a key contributor to Apache Arrow — involved in shaping the project, Positron is backed by deep experience in the tools data scientists already rely on. Just as importantly, it sits within the broader Posit ecosystem, giving it access to an established community and a growing set of tools and resources. For Linux users on Ubuntu or Mint, getting started is straightforward once you know that the terminal installation is the way to go. ## References - [Positron official site](https://positron.posit.co/?ref=datascientyst.com) - [Positron on GitHub](https://github.com/posit-dev/positron?ref=datascientyst.com) - [Positron GitHub Issue #6162 — Linux dependency install issue](https://github.com/posit-dev/positron/issues/6162?ref=datascientyst.com) - [Wes McKinney on GitHub](https://github.com/wesm?ref=datascientyst.com) ### How to Filter a DataFrame for Numeric Values in Pandas URL: https://datascientyst.com/how-to-filter-a-dataframe-for-numeric-values-in-pandas/ Last updated: 2026-08-14T06:25:07.000Z To **filter a DataFrame for numeric values** in Pandas we can: **(1) Use `str.isnumeric()` with boolean indexing** ```python df[df['col'].str.isnumeric()] ``` **(2) Use `pd.to_numeric()` with `errors='coerce'`** ```python df[pd.to_numeric(df['col'], errors='coerce').notna()] ``` **(3) Use regular expressions with `str.match()`** ```python df[df['col'].str.match(r'^\d+$')] ``` --- ### Step 1: Create a DataFrame Assume we have a DataFrame with a string column that contains both numeric and non-numeric values: ```python import pandas as pd data = { 'col': ['123', 'abc', '45', 'xyz', '678', 'hello'] } df = pd.DataFrame(data) ``` DataFrame looks like: | | col | | - | ----- | | 0 | 123 | | 1 | abc | | 2 | 45 | | 3 | xyz | | 4 | 678 | | 5 | hello | --- ### Step 2: Why `df['col'].filter(str.isnumeric)` Fails The original attempt: ```python df['col'].filter(str.isnumeric) ``` does not work because `filter()` is a DataFrame/Series method that **filters labels (index or column names)**, not values. It expects a function that operates on the index labels, not the cell values. Additionally, `str.isnumeric` is a string method, not a callable that `filter()` expects for label-based filtering. --- ### Step 3: Filter Numeric Values with `str.isnumeric()` We can use boolean indexing with the string accessor `str.isnumeric()`: ```python df_numeric = df[df['col'].str.isnumeric()] ``` result: | | col | | - | --- | | 0 | 123 | | 2 | 45 | | 4 | 678 | --- ### Step 4: Filter Numeric Values with `pd.to_numeric()` Another approach is to attempt conversion and keep only successful conversions: ```python df_numeric = df[pd.to_numeric(df['col'], errors='coerce').notna()] ``` result: | | col | | - | --- | | 0 | 123 | | 2 | 45 | | 4 | 678 | --- ### Step 5: Filter Numeric Values with Regular Expressions For more control, use `str.match()` with a regex pattern: ```python df_numeric = df[df['col'].str.match(r'^\d+$')] ``` result: | | col | | - | --- | | 0 | 123 | | 2 | 45 | | 4 | 678 | --- ### Step 6: Convert Filtered Results to Numeric Type After filtering, you may want to convert the remaining values to integers: ```python df_numeric['col'] = df_numeric['col'].astype(int) ``` result: | | col | | - | --- | | 0 | 123 | | 2 | 45 | | 4 | 678 | --- ### Summary | Method | Use Case | | -------------------- | ------------------------------------------------- | | str.isnumeric() | Simple filtering of digit-only strings | | pd.to\_numeric() | Handles mixed types, detects any parseable number | | str.match(r'^\\d+$') | Regex control for custom numeric patterns | The key mistake in the original code was confusing `filter()` (a label-based method) with boolean indexing (a value-based approach). For filtering DataFrame values, always use boolean indexing with `df[condition]`. ### Exploring Economic Data with FRED: A Powerful Source for Data Science Projects URL: https://datascientyst.com/exploring-economic-data-with-fred-a-powerful-source-for-data-science-projects/ Last updated: 2026-07-22T22:07:22.000Z **FRED** (Federal Reserve Economic Data) is a free database run by the Federal Reserve Bank of St. Louis. It holds hundreds of thousands of economic time series from national, international, public, and private sources, plus tools to chart, customize, and export the data. No account needed to browse. 🔗 **Main site:** [fred.stlouisfed.org](https://fred.stlouisfed.org/?ref=datascientyst.com) ## Quick Reference: Where to Go for What | I want to... | Go here | | ---------------------------------- | --------------------------------------------------------------------------------------------------------------- | | Learn what FRED is / its history | [What is FRED?](https://fredhelp.stlouisfed.org/fred/about/about-fred/what-is-fred/?ref=datascientyst.com) | | Browse data by topic (e.g. income) | [Tags: income](https://fred.stlouisfed.org/tags/series?t=income&ref=datascientyst.com) | | Look up a specific dataset | [Example: Real Median Household Income](https://fred.stlouisfed.org/series/MEHOINUSA672N?ref=datascientyst.com) | | See when the next report drops | [Release Calendar](https://fred.stlouisfed.org/releases/calendar?ref=datascientyst.com) | ## The Basics, Fast - **What it is:** A free repository of economic data — GDP, inflation, employment, income, interest rates, and more — maintained since the early 1990s. - **How to find data:** Search by keyword, or browse by source, release, category, or tag. Tags (like [income](https://fred.stlouisfed.org/tags/series?t=income&ref=datascientyst.com)) are the fastest way to explore a topic. - **Every series has its own page**, like [MEHOINUSA672N](https://fred.stlouisfed.org/series/MEHOINUSA672N?ref=datascientyst.com) (Real Median Household Income). Each one includes a chart, download options (CSV, Excel, API), and notes explaining the data. - **New data has a schedule.** Government releases (jobs report, CPI, GDP) come out on fixed dates — check the [release calendar](https://fred.stlouisfed.org/releases/calendar?ref=datascientyst.com) so you're never caught off guard. - **Customize as you go:** switch units (e.g., levels → percent change), change frequency (monthly → annual), or overlay multiple series on one chart. ## Who It's For Students, journalists, investors, researchers, or anyone who wants primary-source economic data instead of a secondhand summary. 👉 Bookmark it: [fred.stlouisfed.org](https://fred.stlouisfed.org/?ref=datascientyst.com) ### A Beginner's Data Science Project: Olympic Medal Winners' Age by Sport URL: https://datascientyst.com/a-beginners-data-science-project-olympic-medal-winners-age-by-sport/ Last updated: 2026-05-01T12:59:59.000Z _This post is for subscribers only._ ### How to Get the Last Column After `str.split()` in a Pandas DataFrame URL: https://datascientyst.com/how-to-get-the-last-column-after-str-split-in-a-pandas-dataframe/ Last updated: 2026-02-17T13:48:16.000Z In this short post we will see how to **split a column into multiple parts** and then extract only the **last component or the last not null value**. This is common with file paths, URLs, codes, delimiter-separated strings and dirty data. For example, if your column contains: - `"user/home/file.txt"`, you might want `"file.txt"`. - `New York,US` \- `US` etc Let's learn how to split a Pandas column and get the last part using simple, efficient methods. We can use the following syntax to margin on a single axis column or row in Pandas: **(1) Use `.str.split()` and `.str[-1]`** ```python df['filename'] = df['path'].str.split('/').str[-1] ``` **(2) Use `rsplit()` with `n=1`** ```python df['filename'] = df['path'].str.rsplit('/', n=1).str[1] ``` **(3) 3: Get next to the last** ```python df['filename'] = df['path'].str.rsplit('/', n=2).str[1] ``` ## Example DataFrame Let’s suppose you have the following DataFrame: ```python import pandas as pd df = pd.DataFrame({ 'path': [ 'home/user/data.csv', 'var/log/errors.log', 'tmp/cache/file.txt' ] }) print(df) ``` Output: | | path | | - | ------------------ | | 0 | home/user/data.csv | | 1 | var/log/errors.log | | 2 | tmp/cache/file.txt | ## 1: Use `.str.split()` and `.str[-1]` The easiest way to extract the **last element** after splitting is to use `str.split()` followed by `.str[-1]`: ```python df['filename'] = df['path'].str.split('/').str[-1] ``` Result: | | path | filename | | - | ------------------ | ---------- | | 0 | home/user/data.csv | data.csv | | 1 | var/log/errors.log | errors.log | | 2 | tmp/cache/file.txt | file.txt | Here, `'/'` is the separator, and `.str[-1]` selects the last item. ## 2: Use `rsplit()` with `n=1` If you want better performance (especially on long strings), you can use `.str.rsplit()` with a limit: ```python df['filename'] = df['path'].str.rsplit('/', n=1).str[1] ``` `rsplit()` splits from the right, and `n=1` ensures only one split is performed. | | path | filename | | - | ------------------ | ---------- | | 0 | home/user/data.csv | data.csv | | 1 | var/log/errors.log | errors.log | | 2 | tmp/cache/file.txt | file.txt | ## 3: Get next to the last Finally if you need to get the next to the last one or X after the split we can control the parameter `n=1`: ```python df['filename'] = df['path'].str.rsplit('/', n=2).str[1] ``` `rsplit()` splits from the right, and `n=2` takes the level we need. | | path | filename | | - | ------------------ | -------- | | 0 | home/user/data.csv | user | | 1 | var/log/errors.log | log | | 2 | tmp/cache/file.txt | cache | ## Summary To extract the last portion of a string in a Pandas column: - Use `str.split()` with `.str[-1]` for simple cases. - Use `str.rsplit()` with `n=1` for slightly better performance on long text. This technique is especially useful for handling file paths, URLs, and delimiter-separated codes. ## Resources - Pandas `str.split()` Documentation [https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.split.html](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.split.html?ref=datascientyst.com) - Pandas `str.rsplit()` Documentation [https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.rsplit.html](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.rsplit.html?ref=datascientyst.com) - [Get last "column" after .str.split() operation on column in pandas DataFrame](https://stackoverflow.com/questions/12504976/get-last-column-after-str-split-operation-on-column-in-pandas-dataframe?ref=datascientyst.com) ### How to Compare Pandas DataFrames When NaNs Are Present URL: https://datascientyst.com/how-to-compare-pandas-dataframes-when-nans-are-present/ Last updated: 2026-02-02T16:34:42.000Z When working with **pandas, comparing DataFrames that contain NaN values can be confusing and error prone**. By defaultin Python, **NaN is not equal to NaN in standard element-wise comparisons**, which often leads to unexpected results. ## Sample data: ```python import numpy as np import pandas as pd df1 = pd.DataFrame([[np.nan,1, np.nan, 3],[2, 1, np.nan,3]]) df2 = df1.copy() ``` | | 0 | 1 | 2 | 3 | | - | --- | - | --- | - | | 0 | NaN | 1 | NaN | 3 | | 1 | 2.0 | 1 | NaN | 3 | ## Why NaN breaks equality checks? In NumPy and pandas, NaN represents missing data. According to IEEE standards, NaN is not equal to anything — even another NaN. This affects comparisons like: ```python df1 == df2 ``` result: 0 1 2 3 0 False True False True 1 True True False True Even if both DataFrames have NaN in the same positions, the result will be False for those cells. ## Use DataFrame.equals() for proper comparison If you want to check whether two DataFrames are truly identical, including NaNs in the same locations, use: ```python df1.equals(df2) ``` result: ``` True ``` Key benefits of equals(): - Treats NaNs in the same position as equal - Requires same shape and values - Ignores index/column type differences if values match **This is the recommended way to compare DataFrames for equality.** Normalize missing values before comparing If one DataFrame uses empty strings and the other uses NaN, equals() will return False. You can standardize them first: ## Replace NaN with empty strings ```python df1.fillna('') == df2.fillna('') ``` result: ``` 0 1 2 3 0 True True True True 1 True True True True ``` Element-wise comparison with NaNs treated as equal If you need element-wise comparison logic, use NumPy: ```python import numpy as np np.isclose(df1, df2, equal_nan=True) ``` ``` [[ True True True True] [ True True True True]] ``` Or for mixed dtypes: ```python (df1.fillna('##NA##') == df2.fillna('##NA##')) ``` ## Summary - NaN != NaN in standard comparisons - Use df.equals() for full DataFrame equality - Normalize NaN and empty values if sources differ - Use NumPy for element-wise comparisons with equal\_nan This avoids false mismatches and ensures consistent DataFrame comparisons when missing data is involved. ## Inspiration The article was inspired from the Kaggle course for data cleaning: - [https://www.kaggle.com/learn/data-cleaning](https://www.kaggle.com/learn/data-cleaning?ref=datascientyst.com) - Lessons - [https://www.kaggle.com/code/alexisbcook/character-encodings](https://www.kaggle.com/code/alexisbcook/character-encodings?ref=datascientyst.com) The homework for deepening the understanding with a dataset of fatal police shootings in the US: - [https://www.kaggle.com/kernels/fork/10824401](https://www.kaggle.com/kernels/fork/10824401?ref=datascientyst.com) Where we have: ```python (police_killings1 != police_killings).melt()['value'].sum() ``` results into: ``` 346 ``` while: ```python (police_killings1.fillna(0) != police_killings.fillna(0)).melt()['value'].sum() ``` results into: ``` 0 ``` And the reason for this are the missing values: ```python for col in police_killings1.columns: if sum(police_killings1[col] != police_killings[col]) > 0: print(col) print(police_killings1[police_killings1[col] != police_killings[col]][col]) print(police_killings[police_killings1[col] != police_killings[col]][col]) ``` result: ``` armed 615 NaN 1551 NaN 1715 NaN 1732 NaN 1825 NaN ... 2487 NaN Name: armed, dtype: object ``` ## Resources - [Pandas DataFrames with NaNs equality comparison](https://stackoverflow.com/questions/19322506/pandas-dataframes-with-nans-equality-comparison?ref=datascientyst.com) ### How to Extract Capital Words from a Pandas DataFrame URL: https://datascientyst.com/how-to-extract-capital-words-from-a-pandas-dataframe/ Last updated: 2026-01-27T13:20:09.000Z If you're working with textual data in a Pandas DataFrame and want to **find all words written in uppercase**, there are several simple ways to do it using Python. A "capital word" here means a word where *every letter is uppercase* (like `JAVA` or `PYTHON`). This is useful when cleaning data, detecting acronyms, or filtering for entries that stand out in text data. ## Example DataFrame Here’s a sample DataFrame we’ll work with: ```python import pandas as pd data = { 'country': ['JAPAN', 'INDIA', 'CHINA', 'FRANCE'], 'users': [30, 15, 25, 3], 'city': ['TOKYO', 'Delhi', 'Beijing', 'PARIS'] } df = pd.DataFrame(data) df ``` This results in: | | country | users | city | | - | ------- | ----- | ------- | | 0 | JAPAN | 30 | TOKYO | | 1 | INDIA | 15 | Delhi | | 2 | CHINA | 25 | Beijing | | 3 | FRANCE | 3 | PARIS | ## 1\. Use `str.isupper()` for Simple Matching The easiest way is to convert each element to a string and check if it is uppercase: ```python import pandas as pd caps = [] for col in df.columns: for val in df[col]: if str(val).isupper(): caps.append(val) print(caps) ``` **Output:** ``` ['JAPAN', 'INDIA', 'CHINA', 'FRANCE', 'TOKYO', 'PARIS'] ``` This method checks if the string version of each cell is all uppercase. ## 2\. Use a Regex Pattern If you prefer regular expressions, you can match uppercase words using a pattern like `r'^[A-Z]+$'`: ```python import re caps = [] for col in df.columns: for val in df[col]: if re.match(r'^[A-Z]+$', str(val)): caps.append(val) print(caps) ``` This matches values made only of uppercase letters and excludes numbers or mixed-case strings. ``` ['JAPAN', 'INDIA', 'CHINA', 'FRANCE', 'TOKYO', 'PARIS'] ``` ## 3\. Apply Across Entire DataFrame You can also use `map()` to test every cell at once and collect uppercase words: ```python caps = df.map(lambda x: str(x).isupper()) uppercase_values = df[caps].stack().tolist() print(uppercase_values) ``` This returns the same list of uppercase words. ``` ['JAPAN', 'TOKYO', 'INDIA', 'CHINA', 'FRANCE', 'PARIS'] ``` ## 4\. Extract words starting with Capital Letters We can also extract words staring with Capital letters but ending in normal case by regex: ```python import re caps = [] for col in df.columns: for val in df[col]: if re.match(r'^[A-Z][a-z]+$', str(val)): caps.append(val) print(caps) ``` result: ``` ['Delhi', 'Beijing'] ``` ## Resources - [https://pandas.pydata.org/docs/user\_guide/text.html](https://pandas.pydata.org/docs/user%5Fguide/text.html?ref=datascientyst.com) - [https://docs.python.org/3/library/re.html](https://docs.python.org/3/library/re.html?ref=datascientyst.com) ### How to Floor a Date to the First Date of That Month in Pandas URL: https://datascientyst.com/how-to-floor-a-date-to-the-first-date-of-that-month-in-pandas/ Last updated: 2025-12-10T22:37:33.000Z Working with dates in Pandas often requires precise manipulations, such as flooring a date to the beginning of its month. This is particularly useful when aggregating data monthly or standardizing timestamps for reporting. In this short guide, we'll explore several efficient methods to achieve this, drawing from proven Pandas techniques. Here's a quick overview of the most common approaches: **(1) Using Period and Timestamp conversion** ```python df['date'] = df['date'].dt.to_period('M').dt.to_timestamp() ``` **(2) Using Timedelta subtraction** ```python df['date'] = df['date'] - pd.to_timedelta(df['date'].dt.day - 1, unit='D') ``` **(3) Using NumPy datetime64\[M\]** ```python df['date'] = df['date'].values.astype('datetime64[M]') ``` These methods handle various edge cases, like dates already on the first of the month. Let's dive into the details. ## 1: Using Period and Timestamp Conversion The most straightforward and highly recommended method leverages Pandas' `to_period` functionality. This converts the datetime series to monthly periods and back to timestamps, effectively flooring to the month's start. Consider this sample DataFrame: ```python dates = pd.date_range('2020-01-15', periods=10, freq='17D') df = pd.DataFrame({ 'date': dates, 'value': np.random.randn(len(dates)) }) df ``` Output: | | date | value | date\_m1 | date\_m2 | date\_m3 | date\_m4 | date\_m5 | date\_norm | date\_m6 | | - | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | | 0 | 2020-01-15 | \-0.427706 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-15 | | 1 | 2020-02-01 | \-1.802907 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-01-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | | 2 | 2020-02-18 | 1.365420 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-18 | | 3 | 2020-03-06 | \-0.090745 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-06 | | 4 | 2020-03-23 | \-1.927388 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-23 | Now apply the flooring: ```python df.index = df.index.to_period('M').to_timestamp() print(df) ``` Output: | | date | value | date\_m1 | date\_m2 | date\_m3 | date\_m4 | date\_m5 | date\_norm | date\_m6 | | - | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | | 0 | 2020-01-15 | \-0.427706 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-15 | | 1 | 2020-02-01 | \-1.802907 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-01-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | | 2 | 2020-02-18 | 1.365420 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-18 | | 3 | 2020-03-06 | \-0.090745 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-06 | | 4 | 2020-03-23 | \-1.927388 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-23 | This approach is vectorized, efficient, and works seamlessly with DataFrame indices or columns. It's particularly robust for resampled data, avoiding errors like non-fixed frequencies in `floor('M')`. ## 2: Using Timedelta Subtraction Another vectorized option subtracts the appropriate number of days to reach the month's start. This method calculates the offset based on the day of the month. Using the same DataFrame: ```python df['date_m3'] = df['date'] - pd.to_timedelta(df['date'].dt.day - 1, unit='D') ``` Output: | | date | value | date\_m1 | date\_m2 | date\_m3 | date\_m4 | date\_m5 | date\_norm | date\_m6 | | - | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | | 0 | 2020-01-15 | \-0.427706 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-15 | | 1 | 2020-02-01 | \-1.802907 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-01-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | | 2 | 2020-02-18 | 1.365420 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-18 | | 3 | 2020-03-06 | \-0.090745 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-06 | | 4 | 2020-03-23 | \-1.927388 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-23 | **Tip:** Use `.dt.day` if operating on a column (e.g., `df['date'].dt.day`). This handles dates already at month-start without shifting them backward. It's compatible with libraries like Dask for larger datasets. ## 3: Using NumPy datetime64\[M\] For a low-level, no-extra-imports solution, cast the datetime array to NumPy's `datetime64[M]` dtype, which truncates to month-start. ```python df.index = pd.to_datetime(df.index).values.astype('datetime64[M]') print(df) ``` Output: | | date | value | date\_m1 | date\_m2 | date\_m3 | date\_m4 | date\_m5 | date\_norm | date\_m6 | | - | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | | 0 | 2020-01-15 | \-0.427706 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-01 | 2020-01-15 | | 1 | 2020-02-01 | \-1.802907 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-01-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | | 2 | 2020-02-18 | 1.365420 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-01 | 2020-02-18 | | 3 | 2020-03-06 | \-0.090745 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-06 | | 4 | 2020-03-23 | \-1.927388 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-01 | 2020-03-23 | This is concise and performant for arrays. If your series is already datetime-typed, skip `pd.to_datetime`. **Note:** In Pandas 1.2+, you can use `df.index.astype('datetime64[M]')` directly. ## 4: Using MonthBegin Offset Pandas offsets provide a clean way to align dates. Subtract a `MonthBegin` offset to floor, but adjust for month-start dates. ```python from pandas.tseries.offsets import MonthBegin df['date_m4'] = df['date'] + pd.offsets.MonthBegin(0) - pd.offsets.MonthBegin(1) df[['date', 'date_m4']] ``` For a more reliable variant combining with the previous DataFrame setup: ```python df.index = df.index + pd.offsets.MonthBegin() - pd.offsets.MonthBegin() ``` This "add and subtract" trick ensures stability for edge cases like January 1st. Output remains the same floored DataFrame. **Tip:** Offsets are great for time-series operations but may require imports. Use with `pd.offsets.MonthBegin(-1)` for flooring if dates aren't at start. ## 5: String Formatting Approach For simplicity (though less efficient due to string conversion), format the date as 'YYYY-MM-01' and parse back. ```python df.index = pd.to_datetime(df.index.dt.strftime('%Y-%m-01')) print(df) ``` Output: ``` value date 1986-01-01 22.93 1986-02-01 15.46 2018-01-01 20.00 2018-02-01 25.00 ``` This works well for quick scripts but avoid in performance-critical code, as string operations aren't vectorized like the above methods. ## Handling Edge Cases and Best Practices - **Already Floored Dates:** Methods 1, 2, and 3 preserve month-starts without shifting. - **Resampled Data:** If your DataFrame comes from `resample('M').sum()`, the Period method (Section 1) integrates best. - **Timezone Awareness:** For tz-aware datetimes, add `.dt.tz_localize('UTC')` post-conversion if needed. - **Performance:** Period conversion is fastest for large Series; test with `%timeit` for your use case. These techniques cover most scenarios for monthly flooring in Pandas. Experiment with the sample code to see what fits your workflow! ## Resources - [Pandas Documentation: Date Offsets](https://pandas.pydata.org/docs/user%5Fguide/timeseries.html?ref=datascientyst.com#date-offsets) - [Pandas Period Handling](https://pandas.pydata.org/docs/reference/api/pandas.Period.html?ref=datascientyst.com) - [Notebook with examples](https://github.com/softhints/Pandas-Tutorials/blob/master/datetime/5.floor-a-date-to-the-first-date-of-that-month-in-pandas.ipynb?ref=datascientyst.com) ### How to Compare Each Value in Pandas Column to All Subsequent Values URL: https://datascientyst.com/how-to-compare-each-value-in-pandas-column-to-all-subsequent-values/ Last updated: 2025-12-09T21:56:26.000Z Learn how to **compare every value in a pandas DataFrame column with all following values** efficiently. ## Sample Data ```python import pandas as pd val = [16, 19, 15, 19, 15] df = pd.DataFrame({'val': val}) ``` | | val | | - | --- | | 0 | 16 | | 1 | 19 | | 2 | 15 | | 3 | 19 | | 4 | 15 | ## 1\. Compare with Subsequent Values Using apply Create a new column with lists of comparison results (e.g., 1 if equal, 0 otherwise) for all later rows: ```python df['match'] = df.apply( lambda row: [ 1 if row['val'] == df.loc[idx, 'val'] else 0 for idx in range(row.name + 1, len(df)) ], axis=1 ) ``` **Result:** | | val | match | | - | --- | -------------- | | 0 | 16 | \[0, 0, 0, 0\] | | 1 | 19 | \[0, 1, 0\] | | 2 | 15 | \[0, 1\] | | 3 | 19 | \[0\] | | 4 | 15 | \[\] | This approach works row-wise and is suitable for moderate-sized DataFrames. ## 2\. Compare Text Values with Subsequent for similarity You need to install library: [python-Levenshtein](https://pypi.org/project/python-Levenshtein/?ref=datascientyst.com) ``` !pip install python-Levenshtein ``` The idea is to match all similarities - i.e. apple and appl: ```python from Levenshtein import ratio df_str = pd.DataFrame({'text': ['apple', 'appl', 'banana', 'apple', 'bananna']}) def is_similar(a, b, threshold=0.8): return 1 if ratio(a, b) >= threshold else 0 df_str['similar_later'] = df_str.apply( lambda row: [ is_similar(row['text'], df_str.loc[idx, 'text']) for idx in range(row.name + 1, len(df_str)) ], axis=1 ) df_str ``` result: | | text | similar\_later | | - | ------- | -------------- | | 0 | apple | \[1, 0, 1, 0\] | | 1 | appl | \[0, 1, 0\] | | 2 | banana | \[0, 1\] | | 3 | apple | \[0\] | | 4 | bananna | \[\] | ## 3\. Compare Values for large DataFrames ```python import numpy as np arr = df['val'].values comparisons = (arr[:, np.newaxis] == arr[np.newaxis, :]) # Full matrix upper_tri = np.triu(comparisons, k=1) ``` result for `array([16, 19, 15, 19, 15])`: ``` array([[False, False, False, False, False], [False, False, False, True, False], [False, False, False, False, True], [False, False, False, False, False], [False, False, False, False, False]]) ``` ## Notes - For large DataFrames, this `apply` method can be slow due to Python loops. - Customize the comparison (e.g., `==` to `>` or a function like Levenshtein distance for strings). - For fully vectorized alternatives, consider NumPy broadcasting if the output format allows (e.g., upper triangular matrix). ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/compare/compare-each-value-in-pandas-column-to-all-subsequent-values.ipynb?ref=datascientyst.com) ### How to Insert Item at Beginning of Pandas Series URL: https://datascientyst.com/how-to-insert-item-at-beginning-of-pandas-series/ Last updated: 2025-12-10T18:18:48.000Z When working with Pandas Series, you may need to add an item at the beginning rather than at the end. While Pandas doesn't have a built-in prepend method, there are several effective ways to accomplish this task. In this short guide, you'll see how to insert an item at the beginning of a Pandas Series. Here you can find the short answer: **(1) Using pd.concat() (Recommended)** ```python pd.concat([pd.Series([1]), a]) ``` **(2) Using list concatenation** ```python pd.Series([1] + a.tolist()) ``` **(3) Using insert with index** ```python a.loc[-1] = 1 a = a.sort_index().reset_index(drop=True) ``` Let's see several useful examples on how to insert an item at the beginning of a Pandas Series. Suppose you have a Series like: ```python import pandas as pd a = pd.Series([2, 3, 4]) print(a) ``` Output: ``` 0 2 1 3 2 4 dtype: int64 ``` ## 1: Insert at beginning using pd.concat() The most straightforward and recommended way to insert an item at the beginning is using `pd.concat()`: ```python import pandas as pd a = pd.Series([2, 3, 4]) result = pd.concat([pd.Series([1]), a]) print(result) ``` Result: ``` 0 1 0 2 1 3 2 4 dtype: int64 ``` Notice the duplicate index values. To reset the index, use `ignore_index=True`: ```python result = pd.concat([pd.Series([1]), a], ignore_index=True) print(result) ``` Result: ``` 0 1 1 2 2 3 3 4 dtype: int64 ``` ## 2: Insert at beginning with custom index If you want to specify a custom index for the new item: ```python import pandas as pd a = pd.Series([2, 3, 4], index=[1, 2, 3]) new_item = pd.Series([1], index=[0]) result = pd.concat([new_item, a]) print(result) ``` Result: ``` 0 1 1 2 2 3 3 4 dtype: int64 ``` This approach maintains the index structure and ensures proper ordering. ## 3: Insert multiple items at the beginning You can insert multiple items at once by creating a Series with multiple values: ```python import pandas as pd a = pd.Series([4, 5, 6]) new_items = pd.Series([1, 2, 3]) result = pd.concat([new_items, a], ignore_index=True) print(result) ``` Result: ``` 0 1 1 2 2 3 3 4 4 5 5 6 dtype: int64 ``` ## 4: Using list concatenation (alternative method) Another approach is converting to a list, adding the item, and converting back: ```python import pandas as pd a = pd.Series([2, 3, 4]) result = pd.Series([1] + a.tolist()) print(result) ``` Result: ``` 0 1 1 2 2 3 3 4 dtype: int64 ``` This method is simple but may be slower for large Series since it involves conversion to and from lists. ## 5: Insert with specific index value If you want to add an item with a specific index that comes before existing indices: ```python import pandas as pd a = pd.Series([2, 3, 4], index=[10, 20, 30]) # Add item with index 0 a.loc[0] = 1 # Sort by index to place it first result = a.sort_index() print(result) ``` Result: ``` 0 1 10 2 20 3 30 4 dtype: int64 ``` ## 6: Insert at beginning preserving data types When inserting items, ensure the data type remains consistent: ```python import pandas as pd # Integer Series a = pd.Series([2, 3, 4], dtype=int) result = pd.concat([pd.Series([1], dtype=int), a], ignore_index=True) print(result) print(f"Data type: {result.dtype}") ``` Result: ``` 0 1 1 2 2 3 3 4 dtype: int64 Data type: int64 ``` If you mix types, Pandas will upcast to a compatible type: ```python import pandas as pd a = pd.Series([2, 3, 4], dtype=int) result = pd.concat([pd.Series([1.5]), a], ignore_index=True) print(result) print(f"Data type: {result.dtype}") ``` Result: ``` 0 1.5 1 2.0 2 3.0 3 4.0 dtype: float64 Data type: float64 ``` ## 7: Performance consideration: Building Series incrementally If you need to add multiple items one by one, it's more efficient to collect them in a list first: ```python import pandas as pd # Less efficient: multiple concatenations a = pd.Series([4, 5]) for value in [3, 2, 1]: a = pd.concat([pd.Series([value]), a], ignore_index=True) print("Result from multiple concatenations:") print(a) # More efficient: build list then create Series values = [1, 2, 3] a = pd.Series([4, 5]) result = pd.Series(values + a.tolist()) print("\nResult from list approach:") print(result) ``` Both produce the same result, but the second approach is significantly faster for large datasets. ## Why there's no prepend() method You might wonder why Pandas doesn't have a built-in `prepend()` method. This is because Series are built on NumPy arrays, where inserting at the beginning requires shifting all existing elements, making it an expensive operation. The design encourages appending (which is more efficient) or using `concat()` for combining Series. ## Common pitfall: Using deprecated append() **Note:** The `append()` method has been deprecated since Pandas 1.4.0 and removed in Pandas 2.0.0\. If you see older code using: ```python # DEPRECATED - Don't use this a.append(pd.Series([1])) ``` Replace it with `pd.concat()`: ```python # Use this instead pd.concat([pd.Series([1]), a], ignore_index=True) ``` ## Summary table: Methods comparison | Method | Pros | Cons | Best For | | ------------------ | -------------------------------- | ----------------------- | -------------------- | | pd.concat() | Clean, flexible, handles indices | Slightly verbose | Most use cases | | List conversion | Simple, readable | Slower for large Series | Small Series | | .loc\[\] with sort | Control over index | Requires sorting | Specific index needs | ## Resources - [Insert Item at Beginning of Pandas Series](https://github.com/softhints/Pandas-Tutorials/blob/master/series/series-insert-beginning-notebook.ipynb?ref=datascientyst.com) - [pandas.concat() documentation](https://pandas.pydata.org/docs/reference/api/pandas.concat.html?ref=datascientyst.com) - [pandas.Series documentation](https://pandas.pydata.org/docs/reference/api/pandas.Series.html?ref=datascientyst.com) - [Working with Pandas Series](https://pandas.pydata.org/docs/user%5Fguide/dsintro.html?ref=datascientyst.com#series) - [Merge, Join, and Concatenate Guide](https://pandas.pydata.org/docs/user%5Fguide/merging.html?ref=datascientyst.com) ### How to Validate Domain Name in Pandas and Python URL: https://datascientyst.com/how-to-validate-domain-name-in-pandas-python/ Last updated: 2025-12-06T11:34:53.000Z When working with user input, web scraping, or data validation, you often need to verify whether a string represents a valid domain name. Python offers several approaches to accomplish this task, from simple regex patterns to specialized libraries. In this guide, you'll learn different methods to validate domain names in Python with practical examples. Here's a quick overview of the solutions: **(1) Using the validators library** ```python import validators validators.domain('example.com') ``` **(2) Using regex patterns** ```python import re pattern = r'^(?:[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?\.)+[a-zA-Z]{2,}$' re.match(pattern, 'example.com') ``` **(3) Using DNS lookup with socket** ```python import socket socket.gethostbyname('example.com') ``` **(4) Using custom validation function** ```python def is_valid_hostname(hostname): if len(hostname) > 255: return False allowed = re.compile(r"(?!-)[A-Z\d-]{1,63}(? 255: return False # Remove trailing dot if present if hostname[-1] == ".": hostname = hostname[:-1] # Check each label allowed = re.compile(r"(?!-)[A-Z\d-]{1,63}(?...") ``` which will raise warning and in future error: `FutureWarning: Passing literal html to 'read_html' is deprecated and will be removed in a future version` You should now use: ```python from io import StringIO pd.read_html(StringIO("...
")) ``` Or for HTML files: ```python pd.read_html("path/to/file.html") # This is still valid ``` ## Use StringIO for HTML strings ```python import pandas as pd from io import StringIO html_data = """
nameage
Alice25
Bob30
""" df = pd.read_html(StringIO(html_data))[0] print(df) ``` result: ``` name age 0 Alice 25 1 Bob 30 ``` ## Use requests for web content ```python import requests import pandas as pd response = requests.get('https://example.com/data.html') df = pd.read_html(StringIO(response.text)) ``` ## Why This Matters Making this change now will: 1. Future-proof your code 2. Remove annoying warning messages 3. Ensure compatibility with upcoming pandas versions **Security concerns** top the list, as accepting arbitrary HTML strings can potentially expose applications to security vulnerabilities. By requiring explicit file paths or URLs, pandas encourages safer data handling practices. **API clarity** is another driving factor. Having a single function that accepts multiple input types can lead to confusion about expected behavior and error handling. Separating these concerns makes the API more predictable and easier to maintain. ### How to Convert a MultiIndex to type String or List of Strings in Pandas URL: https://datascientyst.com/how-to-convert-a-multiindex-to-type-string-or-list-of-strings-in-pandas/ Last updated: 2025-05-16T13:39:45.000Z To convert Pandas MultiIndex to list of strins or Sting we have several options: **(1) Lambda and custom format** ```python midx.to_series().apply(lambda x: '{0}-{1}-{1}'.format(*x)).values ``` **(2) List comprehension** ```python array(['11-21-21', '11-22-22', '12-21-21', '12-22-22'], dtype=object) ``` ## Data | | | | Grade | | -- | -- | -- | ----- | | 11 | 21 | 31 | A | | 22 | 32 | B | | | 12 | 21 | 33 | A | | 22 | 34 | C | | Sample data: ```python import pandas as pd df = pd.DataFrame( {"Grade": ["A", "B", "A", "C"]}, index=[ ["11", "11", "12", "12"], ["21", "22", "21", "22"], ["31", "32", "33", "34"] ] ) ``` ## Lambda and custom format ```python midx.to_series().apply(lambda x: '{0}-{1}-{1}'.format(*x)).values ``` result: ``` array(['11-21-21', '11-22-22', '12-21-21', '12-22-22'], dtype=object) ``` ## List comprehension ```python array(['11-21-21', '11-22-22', '12-21-21', '12-22-22'], dtype=object) ``` result: ``` ['31 21 11', '32 22 11', '33 21 12', '34 22 12'] ``` ## Flatten MultiIndex ```python from itertools import starmap def flat2(midx, sep=''): fstr = sep.join(['{}'] * midx.nlevels) return pd.Index(starmap(fstr.format, midx)) flat2(midx, sep='_') ``` Result: ``` Index(['11_21_31', '11_22_32', '12_21_33', '12_22_34'], dtype='object') ``` For more examples and details on MultiIndex flattening please check: [How to Flatten a MultiIndex in Pandas](https://datascientyst.com/flatten-multiindex-in-pandas/) ## Resource - [How do I convert a MultiIndex to type string](https://stackoverflow.com/questions/39111347/how-do-i-convert-a-multiindex-to-type-string?ref=datascientyst.com) ### Convert Pandas MultiIndex values to New Type -String, Int URL: https://datascientyst.com/convert-pandas-multiindex-values-to-new-type-string-int/ Last updated: 2025-05-16T13:22:27.000Z Here are a few common ways to convert Pandas `MultiIndex` values to string or other data type: **(1) Custom conversion per level** ```python df.index.set_levels(midx.levels[0].astype(int), level=0) \ .set_levels(midx.levels[1].astype(str), level=1) \ .set_levels(midx.levels[2].astype(int), level=2) ``` **(2) Conversion to single type** ```python for level in range(0, len(midx)-1): df.index = df.index.set_levels(midx.levels[level].astype(int), level=level) ``` ## Example MultiIndex ```python import pandas as pd df = pd.DataFrame( {"Grade": ["A", "B", "A", "C"]}, index=[ ["11", "11", "12", "12"], ["21", "22", "21", "22"], ["31", "32", "33", "34"] ] ) ``` data looks like: | | | | Grade | | -- | -- | -- | ----- | | 11 | 21 | 31 | A | | 22 | 32 | B | | | 12 | 21 | 33 | A | | 22 | 34 | C | | ## 1: Iterate each MultiIndex Level and convert We can do custom conversion per level by providing the conversion details explicitly: ```python df.index = df.index.set_levels(midx.levels[0].astype(int), level=0) \ .set_levels(midx.levels[1].astype(float), level=1) \ .set_levels(midx.levels[2].astype(int), level=2) ``` Result: ``` MultiIndex([(11, 21.0, 31), (11, 22.0, 32), (12, 21.0, 33), (12, 22.0, 34)], ) ``` ## 2: Automatic conversion of MI to single type To convert all levels of MultiIndex to a single data types we can use combination of `set_levels` and iteration over each level: ```python for level in range(0, len(midx)-1): df.index = df.index.set_levels(midx.levels[level].astype(int), level=level) ``` ## Resources - [Converting a pandas dataframe with a string type, three level MultiIndex into numeric type objects](https://stackoverflow.com/questions/38620850/converting-a-pandas-dataframe-with-a-string-type-three-level-multiindex-into-nu?ref=datascientyst.com) - [pandas.MultiIndex.set\_levels](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.MultiIndex.set%5Flevels.html?ref=datascientyst.com) ### Pandas: Contains Using Case Insensitive Search URL: https://datascientyst.com/pandas-contains-using-case-insensitive-search/ Last updated: 2025-05-16T12:56:04.000Z To perform **case-insensitive string matching in Pandas**, you can use the `.str` accessor along with regular expressions and the `case=False` parameter **(1) parameter case of str.contains** ```python df1['col'].str.contains("MaX", na=False, case=False) ``` **(2) Margin only on columns** ```python df.query("City.str.lower() == 'new york'") ``` **(3) Margin only on columns** ```python df[df['col'].str.strip().str.match('MaX'.strip(), case=False)] ``` ## 1\. Check if a string contains a word (case-insensitive) ```python import pandas as pd df = pd.DataFrame({ 'name': ['maximum', 'Maxxy', 'MAXa', 'Mini', 'MInimum', 'MaX'] }) # Filter rows where 'name' contains 'bob', ignoring case filtered = df[df['name'].str.contains('max', case=False)] print(filtered) ``` **Output:** ``` name 0 maximum 1 Maxxy 2 MAXa 5 MaX ``` ## 2\. Exact match ignoring case For exact matches (not substrings), normalize strings using `.str.lower()` or `.str.upper()`: ```python df[df['name'].str.lower() == 'max'] ``` Result: ``` name 5 MaX ``` ## 3\. Case-insensitive query We can perform case-insensitive search with Pandas query like: ```python df.query("name.str.lower() == 'max'") ``` Result: ``` name 5 MaX ``` ## Notes: - `case=False` works only with `.str.contains()` and `.str.match()` using regex. - For non-regex comparisons, normalize both sides with `.str.lower()` or `.str.upper()`. ## Summary To make your string filters case-insensitive in Pandas: - Use `.str.contains(..., case=False)` for substring matching. - Use `.str.lower()` or `.str.upper()` for exact value comparisons. ## Resources - [pandas "case insensitive" in a string or "case ignore"](https://stackoverflow.com/questions/51026771/pandas-case-insensitive-in-a-string-or-case-ignore?ref=datascientyst.com) - [pandas.Series.str.contains](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.contains.html?ref=datascientyst.com#pandas-series-str-contains) - [pandas.Series.str.match](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.match.html?ref=datascientyst.com#pandas-series-str-match) ### How to Insert a Row at Top of Pandas DataFrame URL: https://datascientyst.com/how-to-insert-a-row-at-top-of-pandas-dataframe/ Last updated: 2025-04-25T13:59:55.000Z To insert a row at the top or a specific index on DataFrame you can achieve it bt using slicing or `concat`: **(1) Using `pd.concat()` with a list** ```python vals = [1, 2] pd.concat([pd.DataFrame([vals], columns=df.columns), df], ignore_index=True) ``` **(2) Using `df.loc` with manual index shift and sort** ```python df.loc[-1] = [1,2] df.index = df.index + 1 df = df.sort_index() ``` ## 1: Insert a Row on top of Pandas DataFrame Let's say you have a simple DataFrame: ```python import pandas as pd df = pd.DataFrame({ 'name': ['Alice', 'Bob', 'Charlie'], 'age': [25, 30, 35] }) ``` If you want to insert a new row hen you can use the following syntax: ```python vals = ['David', 28] pd.concat([pd.DataFrame([vals], columns=df.columns), df], ignore_index=True) ``` Result: | | name | age | | - | ------- | --- | | 0 | David | 28 | | 1 | Alice | 25 | | 2 | Bob | 30 | | 3 | Charlie | 35 | ### Insert at Position (e.g., Index 1) ```python new_row = pd.DataFrame({'name': ['David'], 'age': [28]}) pd.concat([df.iloc[:1], new_row, df.iloc[1:]]).reset_index(drop=True) ``` Output: | | name | age | | - | ------- | --- | | 0 | David | 28 | | 1 | Alice | 25 | | 2 | Bob | 30 | | 3 | Charlie | 35 | ### Tips - Use `reset_index(drop=True)` after insertion to maintain continuous indexing. - For appending to the end, use: - `df.loc[len(df)] = vals` or - `df = pd.concat([df, new_row])` ## Resource - [Pandas concat documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.concat.html?ref=datascientyst.com) - [DataFrame.loc for assigning values](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.loc.html?ref=datascientyst.com) - [Insert a row to pandas dataframe](https://stackoverflow.com/questions/24284342/insert-a-row-to-pandas-dataframe?ref=datascientyst.com) - [How to concatenate multiple column values into a single column in Pandas dataframe](https://stackoverflow.com/questions/39291499/how-to-concatenate-multiple-column-values-into-a-single-column-in-pandas-datafra?ref=datascientyst.com) ### How to Read CSV Directly from a URL in Pandas and Requests URL: https://datascientyst.com/how-to-read-csv-directly-from-a-url-in-pandas-and-requests/ Last updated: 2025-04-14T13:11:32.000Z Pandas can read CSV files directly from a URL by passing the URL to the `read_csv()` method. This is useful when working with datasets hosted online and for ad hoc tests. We can use the following syntax to **read CSV from URL in Pandas**: **(1) Margin only on rows** ```python import pandas as pd url = "https://raw.githubusercontent.com/softhints/Pandas-Exercises-Projects/refs/heads/main/data/europe_pop.csv" df = pd.read_csv(url) ``` ## Basic Usage You can use `pd.read_csv()` with a URL just like you would with a local file: ```python import pandas as pd url = "https://raw.githubusercontent.com/softhints/Pandas-Exercises-Projects/refs/heads/main/data/europe_pop.csv" df = pd.read_csv(url) ``` This will download and read the CSV file into a DataFrame. ## Notes - The URL must point directly to a `.csv` file - Works with: - `http` - `https` - `ftp` - If the CSV is encoded differently (e.g. UTF-16), you can specify it as param: ```python df = pd.read_csv(url, encoding='utf-16') ``` ## Read Data with Requests In some cases you may need to use Python library requests to read the data first and then load it as DataFrame. This could be related to authentication or security. In this case we can use the following code to read the data: ```python import pandas as pd import io import requests url = "https://raw.githubusercontent.com/softhints/Pandas-Exercises-Projects/refs/heads/main/data/europe_pop.csv" content = requests.get(url).content df = pd.read_csv(io.StringIO(content.decode('utf-8'))) df ``` This approach is ideal for: - quick prototyping - accessing public datasets - integration with APIs or - GitHub-hosted data. ### Further Reading - [Convert text data from requests object to dataframe with pandas](https://stackoverflow.com/questions/39213597/convert-text-data-from-requests-object-to-dataframe-with-pandas?ref=datascientyst.com) - [Pandas read\_csv from url](https://stackoverflow.com/questions/32400867/pandas-read-csv-from-url?ref=datascientyst.com) - [Pandas read\_csv documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) - [Example datasets hosted on GitHub](https://github.com/awesomedata/awesome-public-datasets?ref=datascientyst.com) ### Pandas TypeError 'list' object is not callable - rename Pandas columns URL: https://datascientyst.com/pandas-typeerror-list-object-is-not-callable-rename-pandas-columns/ Last updated: 2025-04-12T12:56:40.000Z The Pandas error `'list' object is not callable` is raised when we try to rename dataframe columns. Usually this means that we try to use list instead of a dict with method: `.rename()`. ```python df.rename(columns=['A', 'B', 'C']) ``` results into: > TypeError: 'list' object is not callable while ```python cols = {'A':'AA', 'B': 'BB', 'C': 'CC'} df.rename(columns=cols) ``` works fine. Here's how to **correctly rename columns in pandas** and avoid the error: ## 1\. Rename with a dictionary using `.rename()` The correct syntax for renaming Pandas columns is: ```python df = df.rename(columns={'old_name': 'new_name'}) ``` Below you can find full example of renaming columns: ```python import pandas as pd df = pd.DataFrame({ "A": [0, 1, 2, 3], "B": [3, 5, 7, 9], "C": [1, 2, 3, 4] }) cols = {'A':'AA', 'B': 'BB', 'C': 'CC'} df.rename(columns=cols) ``` ## 2\. Replace all headers at once using `df.columns = [...]` ```python df.columns = ['AA', 'BB', 'CC'] ``` --- ## 3\. Mistake `TypeError: 'Index' object is not callable` Similar miskate `TypeError: 'Index' object is not callable` is raised when we try to invoke dataframe attribute as a method: ```python df.columns('name', 'age', 'country') ``` This is because `df.columns` is attribe, and we're using `()` as if it were a function. ### Error: ```python df.columns('A', 'B', 'C') ``` ### Solution: ```python df.columns = ['A', 'B', 'C'] ``` This typically happens with incorrect use of parentheses `()` instead of square brackets `[]`. In this short post we saw the reasons and solutions for 2 typical Pandas errors: - `TypeError: 'list' object is not callable` - `TypeError: 'Index' object is not callable` ## Resources - [Comprehensive list of most common pandas errors](https://datascientyst.com/416-pandas-error/) - [Pandas column guides and tutorials](https://datascientyst.com/column/) - [Rename headers - 'list' object is not callable](https://stackoverflow.com/questions/60430241/rename-headers-list-object-is-not-callable?ref=datascientyst.com) ### How to Wrap/Break Long Column Names in Pandas Dataframe URL: https://datascientyst.com/how-to-wrap-break-long-column-names-in-pandas-dataframe/ Last updated: 2025-04-12T11:45:56.000Z To wrap or break long column names in Pandas we can use module `textwrap` and map the column names with new line symbols: **(1) Wrap DataFrame column names** ```python import textwrap cols_wrap = [textwrap.wrap(x, width=20) for x in df.columns] cols_wrap = {' '.join(words) : '
'.join(words) for words in cols_wrap} cols_wrap ``` **(2) Truncate column names** ```python import textwrap cols_wrap = {x: textwrap.wrap(x, width=15)[1] for x in df.columns} ``` This will prevent formatting issues or horizontal overflow on displayed data. Let's see it in more details and examples: ## Data Let's use this data: ```python import pandas as pd df = pd.DataFrame({ "Test Data Type N0 extract 1": [0, 1, 2, 3], "Test Data Type N101 extract 1": [3, 5, 7, 9], "Prod Data Type N0 extract 0": [1, 2, 3, 4], "Prod Data Type N101 extract 0": [0.5, 1.0, 1.5, 2.0], }) ``` which will have long names. If you work with 20+ columns this might be visually hard to digest: | | Test Data Type N0 extract 1 | Test Data Type N101 extract 1 | Prod Data Type N0 extract 0 | Prod Data Type N101 extract 0 | | - | --------------------------- | ----------------------------- | --------------------------- | ----------------------------- | | 0 | 0 | 3 | 1 | 0.5 | | 1 | 1 | 5 | 2 | 1.0 | | 2 | 2 | 7 | 3 | 1.5 | | 3 | 3 | 9 | 4 | 2.0 | ## 1\. Wrap column names We can wrap every column name no matter is it OK or too long, by inserting `
` to break them: ```python import textwrap cols_wrap = [textwrap.wrap(x, width=20) for x in df.columns] cols_wrap = {' '.join(words) : '
'.join(words) for words in cols_wrap} cols_wrap ``` This will create a dictionary: ``` {'Test Data Type N0 extract 1': 'Test Data Type N0
extract 1', 'Test Data Type N101 extract 1': 'Test Data Type N101
extract 1', 'Prod Data Type N0 extract 0': 'Prod Data Type N0
extract 0', 'Prod Data Type N101 extract 0': 'Prod Data Type N101
extract 0'} ``` ```python df.rename(columns=cols_wrap).style.format() ``` Now we can display the DataFrame with wrapped column names: | | Test Data Type N0extract 1 | Test Data Type N101extract 1 | Prod Data Type N0extract 0 | Prod Data Type N101extract 0 | | - | -------------------------- | ---------------------------- | -------------------------- | ---------------------------- | | 0 | 0 | 3 | 1 | 0.500000 | | 1 | 1 | 5 | 2 | 1.000000 | | 2 | 2 | 7 | 3 | 1.500000 | | 3 | 3 | 9 | 4 | 2.000000 | - we can control the lenght of the wrap by - `width=20` - shorter columns will remain the same - the `
` works in Jupyterlab in combination with `.style.format()` - original data is unchanged ## 2\. Truncate Column names We can also break the longer column names by similar approach: ```python import textwrap cols_wrap = {x: textwrap.wrap(x, width=15)[1] for x in df.columns} cols_wrap ``` this time we will have shorter names which consists only from the last part of the wrap: ``` {'Test Data Type N0 extract 1': 'N0 extract 1', 'Test Data Type N101 extract 1': 'N101 extract 1', 'Prod Data Type N0 extract 0': 'N0 extract 0', 'Prod Data Type N101 extract 0': 'N101 extract 0'} ``` result: | | N0 extract 1 | N101 extract 1 | N0 extract 0 | N101 extract 0 | | - | ------------ | -------------- | ------------ | -------------- | | 0 | 0 | 3 | 1 | 0.5 | | 1 | 1 | 5 | 2 | 1.0 | | 2 | 2 | 7 | 3 | 1.5 | | 3 | 3 | 9 | 4 | 2.0 | ## 3\. Transpose for better vertical readability If the dataset is small, sometimes transposing helps: ```python print(df.T.to_string()) ``` or by printing: ```python print(df.T.to_string()) ``` This prints the column names as row labels, which makes even long names easier to read: | | 0 | 1 | 2 | 3 | | ----------------------------- | --- | --- | --- | --- | | Test Data Type N0 extract 1 | 0.0 | 1.0 | 2.0 | 3.0 | | Test Data Type N101 extract 1 | 3.0 | 5.0 | 7.0 | 9.0 | | Prod Data Type N0 extract 0 | 1.0 | 2.0 | 3.0 | 4.0 | | Prod Data Type N101 extract 0 | 0.5 | 1.0 | 1.5 | 2.0 | ## Resource - [Tutorials on Pandas DataFrame columns](https://datascientyst.com/column/) - [Break/wrap long text of column names in Pandas dataframe plain text to\_string output?](https://stackoverflow.com/questions/78129071/break-wrap-long-text-of-column-names-in-pandas-dataframe-plain-text-to-string-ou?ref=datascientyst.com) ### How to Create a Pivot Table and Get Percentages in Pandas URL: https://datascientyst.com/how-to-create-a-pivot-table-and-get-percentages-in-pandas/ Last updated: 2025-04-08T21:05:14.000Z **Pivoting a table and calculating row-wise or column-wise percentages** is a common task in data analysis — often used to understand how values in a row contribute to the row total. Here's how to make a pivot table with it with percentage in **Pandas**: **(1) Calculate row-wise percentage** ```python pivot_pct = pivot.div(pivot.sum(axis=1), axis=0) * 100 ``` **(2) Calculate column-wise percentage** ```python pivot_pct = pivot.div(pivot.sum(axis=0), axis=1) * 100 ``` **(3) Using crosstab and normalize** ```python s=pd.crosstab(index=df['category'],columns=df['type'],values=df['value'], normalize='index',aggfunc='sum').\ add_suffix('_').reset_index() ``` ## Data ```python import pandas as pd # Sample data df = pd.DataFrame({ 'category': ['A', 'A', 'B', 'B', 'B'], 'type': ['X', 'Y', 'X', 'Y', 'Z'], 'value': [75, 25, 20, 50, 30] }) df ``` Original data looks like: | | category | type | value | | - | -------- | ---- | ----- | | 0 | A | X | 75 | | 1 | A | Y | 25 | | 2 | B | X | 20 | | 3 | B | Y | 50 | | 4 | B | Z | 30 | ## 1\. Pivot Table with Row Percentages ```python import pandas as pd # Sample data df = pd.DataFrame({ 'category': ['A', 'A', 'B', 'B'], 'type': ['X', 'Y', 'X', 'Y'], 'value': [10, 30, 20, 80] }) pivot = df.pivot_table(index='category', columns='type', values='value', aggfunc='sum', fill_value=0) pivot_pct = pivot.div(pivot.sum(axis=1), axis=0) * 100 pivot_pct.round(2) ``` **Output:** | type | X | Y | Z | | -------- | --------- | --------- | ----- | | category | | | | | A | 78.947368 | 33.333333 | 0.0 | | B | 21.052632 | 66.666667 | 100.0 | **Explanation** - `.pivot()` groups data like a spreadsheet pivot: it aggregates values based on row and column labels. - `.div(..., axis=0)` divides each row by its sum (row-wise operation). - `* 100` converts proportions to percentages. ### Optional: Format as Percent Strings ```python pivot_pct.map(lambda x: f"{x:.1f}%") ``` ## 2\. Pivot Table with Column Percentages We can normalize pivot table column-wise by: ```python # Pivot the table pivot = df.pivot_table(index='category', columns='type', values='value', aggfunc='sum', fill_value=0) # Calculate row-wise percentage pivot_pct = pivot.div(pivot.sum(axis=0), axis=1) * 100 print(pivot_pct.round(2)) ``` result: | type | X | Y | Z | | -------- | ---- | ---- | ----- | | category | | | | | A | 78.9 | 33.3 | 0.0 | | B | 21.1 | 66.7 | 100.0 | ## 3\. Crosstab and normalize Final way to pivot multiple columns and get the normalized values instead of counts will be by using the `crosstab` method: ```python s=pd.crosstab(index=df['category'],columns=df['type'],values=df['value'], normalize='index',aggfunc='sum').\ add_suffix('_').reset_index() s ``` result: | type | category | X\_ | Y\_ | Z\_ | | ---- | -------- | ---- | ---- | --- | | 0 | A | 0.75 | 0.25 | 0.0 | | 1 | B | 0.20 | 0.50 | 0.3 | ## Resources - [Pandas .pivot\_table() Docs](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.pivot%5Ftable.html?ref=datascientyst.com) - [pandas.pivot](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.pivot.html?ref=datascientyst.com) - [Pandas .div() Docs](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.div.html?ref=datascientyst.com) - [pandas.crosstab](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.crosstab.html?ref=datascientyst.com) - [How can I pivot a table and get the percentage of each row in Python?](https://stackoverflow.com/questions/62067186/how-can-i-pivot-a-table-and-get-the-percentage-of-each-row-in-python?ref=datascientyst.com) This method is great for creating readable summary tables for reports or dashboards. ### How to Format Numbers with Commas for Thousands in Pandas URL: https://datascientyst.com/how-to-format-numbers-with-commas-for-thousands-in-pandas/ Last updated: 2025-04-08T20:53:23.000Z To **display large numbers in a more readable format we can insert commas as thousands separators** in Pandas. This is especially useful when preparing data for presentation or reports. Below is a quick solution to format numbers with commas using Pandas: **(1) Display Only** ```python df.style.format('{:,}') ``` or ```python df.head().style.format("{:,.0f}") ``` **(2) Format parameter for thousands char** ```python df.style.format(thousands=",") ``` **(3) Format multiple columns** ```python col_format = {"sales": "{:,.0f}", "col2": "{:,.0f}"} df.head().style.format(col_format) ``` **(4) pandas format comma thousands** ```python df['sales'].apply(lambda x: f"{x:,}") ``` ## Data Suppose we have the following DataFrame: ```python import pandas as pd df = pd.DataFrame({ 'sales': [1000, 15000, 2500000] }) df ``` data: | | sales | | - | ------- | | 0 | 1000 | | 1 | 15000 | | 2 | 2500000 | ## Display Commas Without Changing Values If you only need the formatted output for display purposes: ```python df.style.format("{:,.0f}") ``` result: | | sales | | - | --------- | | 0 | 1,000 | | 1 | 15,000 | | 2 | 2,500,000 | ## Multiple Columns with Custom Format To apply this to multiple numeric columns: ```python col_format = {"sales": "{:,.0f}", "col2": "{:,.0f}"} df.head().style.format(col_format) ``` result: | | sales | | - | --------- | | 0 | 1,000 | | 1 | 15,000 | | 2 | 2,500,000 | ## Format Column with Commas - New Column Use `.apply` with Python’s built-in `format` function to create a new column or update existing one: ```python import pandas as pd df = pd.DataFrame({ 'sales': [1000, 15000, 2500000] }) df['sales_nice'] = df['sales'].apply(lambda x: f"{x:,}") ``` **Output:** | | sales | sales\_nice | | - | ------- | ----------- | | 0 | 1000 | 1,000 | | 1 | 15000 | 15,000 | | 2 | 2500000 | 2,500,000 | ## Resources - [pandas.io.formats.style.Styler.format](https://pandas.pydata.org/docs/reference/api/pandas.io.formats.style.Styler.format.html?ref=datascientyst.com) - [Convert string, K and M to number, Thousand and Million in Pandas/Python](https://datascientyst.com/convert-string-k-m-to-number-thousand-million-pandas-python/) - [Format a number with commas to separate thousands](https://stackoverflow.com/questions/43102734/format-a-number-with-commas-to-separate-thousands?ref=datascientyst.com) ### Fixing "ValueError: Cannot mix tz-aware with tz-naive values" in Pandas URL: https://datascientyst.com/fixing-valueerror-cannot-mix-tz-aware-with-tz-naive-values-in-pandas/ Last updated: 2025-03-25T16:49:41.000Z When using `pd.to_datetime()` in Pandas, you might encounter the error: ``` ValueError: Cannot mix tz-aware with tz-naive values ``` This happens when: - **timezone-aware** (`tz-aware`) and - **timezone-naive** (`tz-naive`) datetime values exist in the same column. Pandas does not allow this combination for operations like comparison or merging. ## Understanding the Problem - **Timezone-naive** timestamps do not have timezone information. - **Timezone-aware** timestamps include a timezone (`UTC`, `America/New_York`, etc.). Example of mixed values causing the error: ```python import pandas as pd df = pd.DataFrame({"timestamp": ["2024-02-10 12:00:00", "2024-02-10 14:00:00+00:00"]}) df["timestamp"] = pd.to_datetime(df["timestamp"]) ``` which will result into: ``` ValueError: unconverted data remains when parsing with format "%Y-%m-%d %H:%M:%S": "+00:00", at position 1. You might want to try: - passing `format` if your strings have a consistent format; - passing `format='ISO8601'` if your strings are all ISO8601 but not necessarily in exactly the same format; - passing `format='mixed'`, and the format will be inferred for each element individually. You might want to use `dayfirst` alongside this. ``` ## How to Fix the Error ### 1\. Make All Datetimes Timezone-Naive If you don’t need timezones, remove them using `.tz_localize(None)`: ```python df['timestamp'].apply(lambda x: pd.to_datetime(x).tz_localize(None)) ``` ### 2\. Make All Datetimes Timezone-Aware If you need timezones, localize all values explicitly: ```python pd.to_datetime(df['timestamp'], format='mixed', utc=True) ``` or ```python pd.to_datetime(df['timestamp'], format='ISO8601', utc=True) ``` Or specify a different timezone: ```python df["timestamp"] = pd.to_datetime(df["timestamp"]).dt.tz_localize("America/New_York") ``` ## Tips - **Use `.dt.tz_localize(None)`** to remove timezones and make timestamps naive. - **Use `pd.to_datetime(df["col"], utc=True)`** to ensure consistency in timezone-aware timestamps. - **Convert timezones** before merging or comparing datetime values. ## Resources - [Handling CSV with timezone-aware and timezone-naive datetime column](https://stackoverflow.com/questions/68182325/handling-csv-with-timezone-aware-and-timezone-naive-datetime-column?ref=datascientyst.com) - [How to Convert Datetime to the Same Timezone in Pandas DataFrame](https://datascientyst.com/convert-datetime-the-same-timezone-pandas-dataframe/) - [Timezone-aware and naive timestamp in Pandas](https://datascientyst.com/timezone-and-naive-timestamp-in-pandas/) ### How to Split Strings and Extract the N-th Element in a Pandas DataFrame URL: https://datascientyst.com/how-to-split-strings-and-extract-the-n-th-element-in-a-pandas-dataframe/ Last updated: 2025-03-18T05:05:37.000Z Working with text data in Pandas, you may need to split strings based on a n-th occuramce of delimiter and extract specific parts. This is useful for parsing URLs, file paths, or structured data. Pandas provides efficient ways to handle such operations with `str.split()` and `expand=True`. **(1) Split the string at the nth occurrence and keep the first part** ```python df['column_name'].str.split('-', n=2).str[0] ``` **(2) Extract the nth part from a split string** ```python df['column_name'].str.split('-', expand=True)[2] ``` ## 1\. Sample data ```python import pandas as pd data = ['https://example.com/search?q=avatar', 'https://example.com/profile/avatar', 'https://example.com/map'] df = pd.DataFrame({'text': data}) df ``` data looks like: | | text | first\_three | | - | ----------------------------------- | ------------ | | 0 | https://example.com/search?q=avatar | https: | | 1 | https://example.com/profile/avatar | https: | | 2 | https://example.com/map | https: | ## 2\. Splitting a String at the n-th Occurrence To split a string only at the n-th occurrence of a delimiter, use the `n` parameter of `str.split()`. We can extract the last part of the URL by: ```python df['text'].str.split('/', n=3, expand=True)[3] ``` **Output:** ``` 0 search?q=avatar 1 profile/avatar 2 map Name: 3, dtype: object ``` - `n=3` ensures only 3 splits occur - `[3]` extracts the 3rd part of the split below you can find the resulted dataframe from the split: | | 0 | 1 | 2 | 3 | | - | ------ | - | ----------- | --------------- | | 0 | https: | | example.com | search?q=avatar | | 1 | https: | | example.com | profile/avatar | | 2 | https: | | example.com | map | ## 3\. Extracting the nth Element from a Split String If you need to extract the nth part of the split string, use `expand=True` to create multiple columns. ```python df[['protocol', 'empty', 'domain', 'method', 'param']] = df['text'].str.split('/', expand=True) df ``` **Output:** | | text | first\_three | protocol | empty | domain | method | param | | - | ----------------------------------- | ------------ | -------- | ----- | ----------- | --------------- | ------ | | 0 | https://example.com/search?q=avatar | https: | https: | | example.com | search?q=avatar | None | | 1 | https://example.com/profile/avatar | https: | https: | | example.com | profile | avatar | | 2 | https://example.com/map | https: | https: | | example.com | map | None | ## 4\. Keeping Only the Last Two Parts of a Split String For cases like domain extraction (`example.com`, `www.example.com`), keep only the last two parts. Or keep the domain and the method from URL: ```python df['text'].apply(lambda x: '/'.join(x.split('/')[-2:])) ``` **Output:** ``` 0 example.com/search?q=avatar 1 profile/avatar 2 example.com/map Name: text, dtype: object ``` - `x.split('/')[-2:]` keeps only the last two elements. - `'/'.join(...)` reconstructs the truncated string. ## 5\. Conclusion Pandas provides multiple ways to split strings based on the nth occurrence of a delimiter. Whether you need to keep a portion of the string, extract a specific element, or retain only the last few parts, `str.split()` and `apply()` are effective tools for data transformation. --- ### **Resources** - [Pandas str.split() Documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.split.html?ref=datascientyst.com) - [StackOverflow: Splitting String in a Pandas DataFrame](https://stackoverflow.com/questions/14745022/how-to-split-a-column-into-multiple-columns-in-pandas?ref=datascientyst.com) - [Splitting nth elements in a string in a pandas dataframe](https://stackoverflow.com/questions/71764317/splitting-nth-elements-in-a-string-in-a-pandas-dataframe?ref=datascientyst.com) - [How to Strip a String After the Nth Occurrence of a Character in Python](https://softhints.com/how-to-strip-a-string-after-the-nth-occurrence-of-a-character-in-python/?ref=datascientyst.com) ### How to Estimate the Memory Usage of a Pandas DataFrame URL: https://datascientyst.com/how-to-estimate-the-memory-usage-of-a-pandas-dataframe/ Last updated: 2025-03-08T12:31:56.000Z When working with large datasets, it's important to estimate how much memory a **Pandas DataFrame** will consume. This helps optimize performance and prevent memory errors. **(1) Calculate memory usage per column** ```python df.memory_usage() ``` **(2) Interrogating object dtypes for system-level memory consumption** ```python df.memory_usage(deep=True) ``` **(3) Total Memory Usage** ```python print(df.memory_usage(deep=True).sum(), "bytes") ``` ## 1\. Use `df.memory_usage()` Pandas provides the `memory_usage()` method to calculate memory usage per column: ```python import pandas as pd import numpy as np # Sample DataFrame df = pd.DataFrame({ "col1": np.random.randint(0, 100, 1000), "col2": np.random.random(1000), "col3": [str(i) for i in range(1000)] }) # Memory usage per column print(df.memory_usage(deep=True)) # 'deep=True' includes object types ``` result: ``` Index 146309736 count 146309736 date 146309736 hour 146309736 url 146309736 ... dtype: int64 ``` ## 2\. Estimate Total Memory Usage To get the total DataFrame size in bytes: ```python print(df.memory_usage(deep=True).sum(), "bytes") ``` result: ``` 14610252261 bytes ``` ## 3\. How memory\_usage deep=True works To get the total DataFrame size in bytes: ```python df.memory_usage(deep=False) ``` result: ``` Index 146309736 count 146309736 date 146309736 hour 146309736 url 146309736 netloc 146309736 ... dtype: int64 ``` vs ```python df.memory_usage(deep=True) ``` result: ``` Index 146309736 count 146309736 date 146309736 hour 146309736 url 4123904171 netloc 1005879435 ... dtype: int64 ``` You can notice the difference which is impacting the object columns: ``` count int64 date datetime64[ns] hour int64 url object netloc object ... dtype: object ``` ## 4\. Optimizing Memory Usage To reduce memory usage: - Convert **integers** to smaller types (`int8`, `int16`). - Use **categorical types** for repetitive strings. Good indicator for this is testing values with `df.describe(include='all')` - Convert **floating points** to lower precision (`float16`, `float32`). ```python df["col1"] = df["col1"].astype("int16") df["col3"] = df["col3"].astype("category") print(df.memory_usage(deep=True).sum(), "bytes after optimization") ``` after optimizing we can find the difference: ``` Index 139.53 MB count 139.53 MB date 139.53 MB hour 139.53 MB url 3.84 GB netloc 959.28 MB dtype: object ``` vs ``` Index 139.53 MB count 69.77 MB date 139.53 MB hour 17.44 MB url 3.84 GB netloc 17.44 MB dtype: object ``` which is result of applying: ```python df.hour = pd.to_numeric(df.hour, downcast='integer') df['count'] = pd.to_numeric(df['count'], downcast='integer') df['netloc'] = df['netloc'].astype('category') ``` ## 5\. Human readable memory info This answers is inspired by SO shared in the resource section: ```python suffixes = ['B', 'KB', 'MB', 'GB', 'TB', 'PB'] def humansize(nbytes): i = 0 while nbytes >= 1024 and i < len(suffixes)-1: nbytes /= 1024. i += 1 f = ('%.2f' % nbytes).rstrip('0').rstrip('.') return '%s %s' % (f, suffixes[i]) df.memory_usage(index=True, deep=True).apply(humansize) ``` output: ``` Index 139.53 MB count 139.53 MB date 139.53 MB hour 139.53 MB url 3.84 GB netloc 959.28 MB dtype: object ``` ## Resources - [How to estimate how much memory a Pandas' DataFrame will need?](https://stackoverflow.com/questions/18089667/how-to-estimate-how-much-memory-a-pandas-dataframe-will-need?ref=datascientyst.com) - [pandas.DataFrame.memory\_usage](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.memory%5Fusage.html?ref=datascientyst.com) ### How To Find the Closest Values in a Pandas Series to a Number URL: https://datascientyst.com/how-to-find-the-closest-values-in-a-pandas-series-to-a-number/ Last updated: 2025-02-21T14:31:50.000Z In this post we will see how to find the closest values in a pandas Series to a given number. Here you can find two short solutions: **(1) Find the single closest value in a Pandas Series** ```python closest_value = df['column_name'].iloc[(df['column_name'] - input_value).abs().idxmin()] ``` **(2) Find the N closest values** ```python n_closest = df.loc[(df['column_name'] - input_value).abs().nsmallest(N).index, 'column_name'] ``` ![](https://datascientyst.com/content/images/2025/02/how-to-find-the-closest-values-in-a-pandas-series-to-a-number.png) Finding the closest value to a given input in a Pandas Series is useful for rounding numbers, performing nearest-neighbor lookups, or handling continuous data. ## 1\. Sample Data Let's create a sample dataset: ```python import pandas as pd import numpy as np data = {'values': [5, 12, 20, 28, 35, 42]} df = pd.DataFrame(data) df ``` This creates a DataFrame: | | values | | - | ------ | | 0 | 5 | | 1 | 12 | | 2 | 20 | | 3 | 28 | | 4 | 35 | | 5 | 42 | --- ## 2\. Find the Single Closest Value To find the closest value to `25`, we use `idxmin()` on the absolute difference: ```python value = 25 closest_value = df['values'].iloc[(df['values'] - value).abs().idxmin()] closest_value ``` **Output:** ``` 28 ``` The closest value to `25` is `28`. ## 3\. Sort by proximity to value We can sort values in a Series based on their proximity to a given value. Be careful for NaN values which can bring unexpected results: ```python import pandas as pd import numpy as np data = {'values': [ 20, np.NaN, 28, 35, 42, 1, -1, 5, -5, 12,]} df = pd.DataFrame(data) value = 0 df.loc[(df['values']).abs().nsmallest(5).index, 'values'] ``` Here are the 5 closest numbers to the given value: ``` 5 1.0 6 -1.0 7 5.0 8 -5.0 9 12.0 Name: values, dtype: float64 ``` ## 4\. Find the N Closest Values To find the `3` closest values to `25`: ```python n = 3 df.loc[(df['values'] - value).abs().nsmallest(n).index, 'values'] ``` **Output:** ``` 3 28 2 20 4 35 Name: values, dtype: int64 ``` The three closest values to `25` are: - `20` - `28` - `35` ## 5\. Find the Closest Larger or Smaller Value To find the closest larger or the smaller number from Pandas Series to a given number we can use: - **Find the closest larger value:** ```python df.loc[df['values'] >= value, 'values'].min() ``` **Output:** ``` 28 ``` - **Find the closest smaller value:** ```python df.loc[df['values'] <= value, 'values'].max() ``` **Output:** ``` 20 ``` ## 6\. Handling Missing or Empty Data If the Series contains `NaN` values, you should drop them before finding the closest value: ```python df = df.dropna() ``` To handle the case where no valid values exist (e.g., all values are larger/smaller), use: ```python if df['values'].lt(input_value).any(): smallest = df.loc[df['values'] < input_value, 'values'].max() else: smallest = None ``` ## 7\. Conclusion Finding the closest value in a Pandas Series is simple using `.idxmin()` for a single value or `.nsmallest()` for multiple values. You can also filter values to find the closest larger or smaller number as needed. ## **Resources** - [Pandas Series.idxmin() Documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.idxmin.html?ref=datascientyst.com) - [Pandas Series.nsmallest() Documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.nsmallest.html?ref=datascientyst.com) - [How do I find the closest values in a Pandas series to an input number?](https://stackoverflow.com/questions/30112202/how-do-i-find-the-closest-values-in-a-pandas-series-to-an-input-number?ref=datascientyst.com) ### How to Calculate Days Elapsed Since a Certain Date in Pandas URL: https://datascientyst.com/how-to-calculate-days-elapsed-since-a-certain-date-in-pandas/ Last updated: 2025-02-18T13:52:22.000Z In this guide, I'll show you how to calculate days elapsed since a certain date in Pandas. Calculating the number of days elapsed since a certain date is a common task in data analysis. **(1) Calculate days elapsed since today** ```python df['days_elapsed'] = (pd.to_datetime('today') - pd.to_datetime(df['date_column'])).dt.days ``` **(2) Using Specific date** ```python reference_date = pd.to_datetime('2023-12-31') df['days_since_fixed'] = (reference_date - df['date_column']).dt.days ``` ## 1\. Sample Data Let's start with a sample dataset: ```python import pandas as pd data = {'date_column': ['2023-01-01', '2022-06-15', '2024-02-01']} df = pd.DataFrame(data) df['date_column'] = pd.to_datetime(df['date_column']) ``` This will produce a DataFrame like: | | date\_column | | - | ------------ | | 0 | 2023-01-01 | | 1 | 2022-06-15 | | 2 | 2024-02-01 | --- ## 2\. Calculating Days Elapsed - Today To calculate the number of days elapsed since a given date, we subtract the date column from today’s date: ```python df['days_elapsed'] = (pd.to_datetime('today') - df['date_column']).dt.days ``` ### **Example Output** | | date\_column | days\_elapsed | | - | ------------ | ------------- | | 0 | 2023-01-01 | 405 | | 1 | 2022-06-15 | 605 | | 2 | 2024-02-01 | 9 | ## 3\. Handling Missing Values If the column contains missing values (`NaT`), Pandas will return `NaN`. You can replace missing values with a default number: ```python df['days_elapsed'] = df['days_elapsed'].fillna(0) ``` ## 4\. Using a Fixed Date If you need to calculate the days elapsed from a fixed reference date instead of today, specify the reference date: ```python reference_date = pd.to_datetime('2023-12-31') df['days_since_fixed'] = (reference_date - df['date_column']).dt.days ``` ### **Example Output** | | date\_column | days\_since\_fixed | | - | ------------ | ------------------ | | 0 | 2023-01-01 | 364 | | 1 | 2022-06-15 | 564 | | 2 | 2024-02-01 | \-32 | ## 5\. Conclusion Using Pandas, you can quickly compute the number of days that have passed since a certain date or until a future date. These calculations are useful for time-based analytics, forecasting, and tracking events. ## **Resources** - [Pandas to\_datetime documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.to%5Fdatetime.html?ref=datascientyst.com) - [Pandas time series guide](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/timeseries.html?ref=datascientyst.com) ### How to Count Special Characters in a Column in Pandas URL: https://datascientyst.com/how-to-count-special-characters-in-a-column-in-pandas/ Last updated: 2025-02-17T22:33:20.000Z In this guide, I'll show you **how to count special characters in a column using Pandas.** Whether you want to count special characters row-wise or in the entire column or a single column, these methods will help. **(1) Count special characters in each row** ```python df['column_name'].str.count(r'[^a-zA-Z0-9\s]') ``` **(2) Count total special characters in the entire column** ```python df['column_name'].str.count(r'[^a-zA-Z0-9\s]').sum() ``` **(3) Count occurrences of a specific special character (e.g., `@`)** ```python df['column_name'].str.count(r'@').sum() ``` **(4) Count the chars by difference of the lengths** ```python df['text'].str.len() - df['text'].str.replace(r'.', '').str.len() ``` ![](https://datascientyst.com/content/images/2025/02/how-to-count-special-characters-in-a-column-in-pandas.webp) ## 1: Example DataFrame Let's create a sample DataFrame with text values: ```python import pandas as pd data = { 'text': ['Hello@World!', 'Python#Pandas$', 'Data&Science*', 'Special_Chars%'] } df = pd.DataFrame(data) ``` ### **Output:** | | text | | - | --------------- | | 0 | Hello@World! | | 1 | Python#Pandas$ | | 2 | Data&Science\* | | 3 | Special\_Chars% | ## 2: Count Special Characters in Each Row To count special characters in each row, use `.str.count()` with a regex pattern: ```python df['special_char_count'] = df['text'].str.count(r'[^a-zA-Z0-9\s]') ``` ### **Output:** | | text | special\_char\_count | | - | --------------- | -------------------- | | 0 | Hello@World! | 2 | | 1 | Python#Pandas$ | 2 | | 2 | Data&Science\* | 2 | | 3 | Special\_Chars% | 1 | ## 3: Count Total Special Characters in Column To count the total number of special characters across all rows: ```python total_special_chars = df['text'].str.count(r'[^a-zA-Z0-9\s]').sum() ``` ### **Output:** ``` 8 ``` ## 4: Count Occurrences of a Specific Special Character (e.g., `@`) If you need to count how many times a specific character (like `@`) appears in the column: ```python at_count = df['text'].str.count(r'@').sum() at_count = df['text'].str.count(r'@') ``` ### **Output:** ``` 1 ``` and ``` 0 1 1 0 2 0 3 0 ``` ## 5: Count Special Characters Across Multiple Columns If you want to check special characters in multiple text columns: ```python cols = ['text'] # Add more columns if needed df['special_char_count'] = df[cols].apply(lambda x: x.str.count(r'[^a-zA-Z0-9\s]')).sum(axis=1) ``` ## 6: Count number of dots in column As an alternative solution we can remove the characters from the column and get the difference from the original length: ```python df['text'].str.len() - df['text'].str.replace(r'.', '').str.len() ``` ## **Conclusion** This guide covered multiple ways to count special characters in Pandas, including: - Counting special characters in each row - Summing special characters across the entire column - Finding occurrences of specific special characters These methods are useful for **text processing, data cleaning, and validation tasks.** ## **Resources** - [Pandas .str.count() Documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.count.html?ref=datascientyst.com) - [Regular Expressions in Python](https://docs.python.org/3/library/re.html?ref=datascientyst.com) - [Pandas String Operations](https://pandas.pydata.org/docs/user%5Fguide/text.html?ref=datascientyst.com) - [How to count special chars in column in Pandas?](https://stackoverflow.com/questions/59687650/how-to-count-special-chars-in-column-in-pandas?ref=datascientyst.com) ### How to Count the Occurrences of a Specific Value in Pandas DataFrame URL: https://datascientyst.com/how-to-count-the-occurrences-of-a-specific-value-in-pandas-dataframe/ Last updated: 2025-02-17T22:18:29.000Z In this short guide, I'll show you **how to count occurrences of a specific value in Pandas.** Whether you want to count values in a single column or across the whole DataFrame, this guide has you covered. **(1) Count occurrences in a single column** ```python df['column_name'].value_counts().get('target_value', 0) ``` **(2) Count occurrences across the entire DataFrame** ```python (df == 'target_value').sum().sum() ``` **(3) Count occurrences using `.shape`** ```python df[df['column_name'] == 'target_value'].shape[0] ``` **(4) Count occurrences using `len()`** ```python len(df[df['column_name'] == 'target_value']) ``` **(5) Count occurrences using `.query()`** ```python df.query('column_name == "target_value"').column_name.count() ``` **(6) Count occurrences using boolean mask with `.sum()`** ```python (df['column_name'] == 'target_value').sum() ``` **(7) Count occurrences using NumPy array** ```python (df['column_name'].values == 'target_value').sum() ``` ![](https://datascientyst.com/content/images/2025/02/how-to-count-the-occurrences-of-a-specific-value-in-pandas-dataframe.webp) ## 1: Example DataFrame Let's create a sample DataFrame: ```python import pandas as pd data = { 'A': ['apple', 'banana', 'apple', 'orange', 'banana'], 'B': ['apple', 'apple', 'banana', 'banana', 'apple'] } df = pd.DataFrame(data) ``` ### **Output:** | | A | B | | - | ------ | ------ | | 0 | apple | apple | | 1 | banana | apple | | 2 | apple | banana | | 3 | orange | banana | | 4 | banana | apple | ## 2: Count Occurrences in a Column To count occurrences of `"apple"` in column `A`: ```python df['A'].value_counts().get('apple', 0) ``` ### **Output:** ``` 2 ``` ## 3: Count Occurrences Across the Entire DataFrame To count `"apple"` occurrences in all columns: ```python (df == 'apple').sum().sum() ``` ### **Output:** ``` 4 ``` ## 4: Count Multiple Values If you want to count multiple values in a column: ```python df['A'].value_counts() ``` ### **Output:** ``` banana 2 apple 2 orange 1 Name: A, dtype: int64 ``` ## Conclusion In this guide, we covered: - Counting occurrences of a value in a specific column - Counting occurrences across the entire DataFrame - Using `.value_counts()` for frequency analysis ## Resources - [Pandas .value\_counts() Documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.value%5Fcounts.html?ref=datascientyst.com) - [Pandas Boolean Indexing](https://pandas.pydata.org/docs/user%5Fguide/indexing.html?ref=datascientyst.com#boolean-indexing) - [Pandas Summing Boolean Values](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.sum.html?ref=datascientyst.com) - [Python Pandas Counting the Occurrences of a Specific value](https://datascientyst.com/questions/35277075/python-pandas-counting-the-occurrences-of-a-specific-value) ### How To Split Column Data Based on Condition in Pandas URL: https://datascientyst.com/how-to-split-column-data-based-on-condition-in-pandas/ Last updated: 2025-02-17T21:53:01.000Z In this short guide, I'll show you **how to split a column based on condition in Pandas DataFrame.** The example will show how to split a column containing URLs and extract only the last two parts (domain and top-level domain - TLD) **(1) Quick Solution Using `.` as a Separator** ```python df['domain'] = df['url'].apply(lambda x: '.'.join(x.split('.')[-2:])) ``` **(2) Using `rsplit()` for a More Efficient Split** ```python df['domain'] = df['url'].str.rsplit('.', n=2).str[-2:].str.join('.') ``` **(3) Using lamdba for conditional split** ```python def find_value_column(row): if row.url.count('.') == 2: return row['url'].split('.', 1)[1] else: return row['url'] df['domain'] = df.apply(find_value_column, axis=1) ``` ![](https://datascientyst.com/content/images/2025/02/how-to-split-column-data-based-on-condition-in-pandas.webp) ## 1: Example DataFrame with URLs Let's create a DataFrame with URLs containing one or two dots: ```python import pandas as pd # Sample data data = { 'url': ['example.com', 'www.example.com', 'test.org', 'blog.test.org'] } df = pd.DataFrame(data) ``` ### **Output:** | | url | | - | --------------- | | 0 | example.com | | 1 | www.example.com | | 2 | test.org | | 3 | blog.test.org | ## 2: Extracting the Netloc and Domain To keep only the last two parts of the domain, we can use `split('.')` and take the last two elements: ```python df['domain'] = df['url'].apply(lambda x: '.'.join(x.split('.')[-2:])) ``` ### **Output:** | | url | domain | | - | --------------- | ----------- | | 0 | example.com | example.com | | 1 | www.example.com | example.com | | 2 | test.org | test.org | | 3 | blog.test.org | test.org | ## 3: Optimized Solution Using `rsplit()` A more efficient approach is using `rsplit()`, which splits from the right and limits the number of splits: ```python df['domain'] = df['url'].str.rsplit('.', n=2).str[-2:].str.join('.') ``` This method avoids unnecessary splits and is faster for large datasets. ## Conclusion In this guide, we learned how to: - Extract the last two parts of a domain name - Use `split('.')` with `apply()` for flexible extraction - Use `rsplit()` for a more optimized approach ## Resources - [Pandas .str.split() Documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.split.html?ref=datascientyst.com) - [Pandas .apply() Function](https://pandas.pydata.org/docs/reference/api/pandas.Series.apply.html?ref=datascientyst.com) - [Pandas .rsplit() for Right Splitting](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.rsplit.html?ref=datascientyst.com) ### Count Characters in a Column and Create a new Length Column in Pandas URL: https://datascientyst.com/count-characters-in-a-column-and-create-a-new-length-column-in-pandas/ Last updated: 2025-02-17T22:01:40.000Z In this short guide, I'll show you **how to count the number of characters in a string column and store the result as a new column in a Pandas DataFrame**. **(1) Quick and Fast Solution Using `.str.len()`** ```python df['char_count'] = df['text'].str.len() ``` **(2) Using `apply(len)` as an Alternative** ```python df['char_count'] = df['text'].apply(len) ``` ![](https://datascientyst.com/content/images/2025/02/count-characters-in-a-column-and-create-a-new-column-in-pandas.webp) This is useful when analyzing text data, measuring string lengths, or filtering based on character count. For example when you need to filter out some outliers or analyse data based on the groups. ## 1: Example DataFrame with Text Column Let's create a sample DataFrame with a column containing text: ```python import pandas as pd # Sample data data = { 'text': ['Hello', 'Pandas', 'DataFrame', 'Python is fun!'] } df = pd.DataFrame(data) df ``` ### **Output:** | | text | | - | -------------- | | 0 | Hello | | 1 | Pandas | | 2 | DataFrame | | 3 | Python is fun! | ## 2: Count Characters in Each String The easiest way to count characters in a string column is using `.str.len()`: ```python df['char_count'] = df['text'].str.len() ``` ### **Output:** | | text | char\_count | | - | -------------- | ----------- | | 0 | Hello | 5 | | 1 | Pandas | 6 | | 2 | DataFrame | 10 | | 3 | Python is fun! | 14 | ## 3: Alternative Method Using `apply(len)` Another way to count characters is by using `apply(len)`, which applies Python’s built-in `len()` function to each row: ```python df['char_count'] = df['text'].apply(len) ``` This method is slightly slower than `.str.len()` but works well in most cases. ## Conclusion In this guide, we learned how to: - Count the number of characters in a string column - Create a new column to store the character count - Use `.str.len()` for better performance - Use `apply(len)` as an alternative ## **Resources** - [Pandas .str.len() Documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.len.html?ref=datascientyst.com) - [Pandas .apply() Function](https://pandas.pydata.org/docs/reference/api/pandas.Series.apply.html?ref=datascientyst.com) - [Python len() Function](https://docs.python.org/3/library/functions.html?ref=datascientyst.com#len) ### How to Combine Date and Time Columns with Pandas URL: https://datascientyst.com/how-to-combine-date-and-time-columns-with-pandas/ Last updated: 2025-02-14T13:55:22.000Z In this short guide, I'll show you **how to combine separate Date and Time columns into a single DateTime column in Pandas**. When working with datasets, dates and times are often stored separately. Merging them into a single column can help with time-series analysis, sorting, and filtering. **(1) Quick Solution Using `pd.to_datetime()`** ```python df['datetime'] = pd.to_datetime(df['date'] + ' ' + df['time']) ``` **(2) Handling Different Formats (e.g., 12-hour format with AM/PM)** ```python df['datetime'] = pd.to_datetime(df['date'] + ' ' + df['time'], format="%Y-%m-%d %I:%M %p") ``` ## 1: Example of Separate Date and Time Columns Let’s say we have a dataset with two columns: `date` and `time`: ```python import pandas as pd # Sample data data = { 'date': ['2024-02-10', '2024-02-11', '2024-02-12'], 'time': ['12:30:00', '14:45:00', '09:15:00'] } df = pd.DataFrame(data) ``` ### **Output:** | | date | time | | - | ---------- | -------- | | 0 | 2024-02-10 | 12:30:00 | | 1 | 2024-02-11 | 14:45:00 | | 2 | 2024-02-12 | 09:15:00 | ## 2: Combine Date and Time into DateTime Column To merge them, we can use `pd.to_datetime()`: ```python df['datetime'] = pd.to_datetime(df['date'] + ' ' + df['time']) ``` ### **Output:** | | date | time | datetime | | - | ---------- | -------- | ------------------- | | 0 | 2024-02-10 | 12:30:00 | 2024-02-10 12:30:00 | | 1 | 2024-02-11 | 14:45:00 | 2024-02-11 14:45:00 | | 2 | 2024-02-12 | 09:15:00 | 2024-02-12 09:15:00 | The new column `datetime` is now in DateTime format, which allows for easier manipulation and analysis. ![](https://datascientyst.com/content/images/2025/02/how-to-combine-date-and-time-columns-with-pandas.png) ## 3: Handling Different Formats If your dataset has different date or time formats, Pandas can automatically parse them, or you can specify a format: ```python df['datetime'] = pd.to_datetime(df['date'] + ' ' + df['time'], format="%Y-%m-%d %H:%M:%S") ``` For a **12-hour format with AM/PM**, use: ```python df['datetime'] = pd.to_datetime(df['date'] + ' ' + df['time'], format="%Y-%m-%d %I:%M %p") ``` ## Conclusion In this guide, we covered how to: - Merge separate `date` and `time` columns - Convert them into a single DateTime column - Handle different time formats This method is useful when working with time-based datasets, logs, and time-series analysis. ## Resources - [Pandas to\_datetime Documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.to%5Fdatetime.html?ref=datascientyst.com) - [Working with Date and Time in Pandas](https://pandas.pydata.org/docs/user%5Fguide/timeseries.html?ref=datascientyst.com) - [Python strftime and strptime Format Codes](https://docs.python.org/3/library/datetime.html?ref=datascientyst.com#strftime-strptime-behavior) ### How to Compare Two Excel or CSV Files Using Pandas URL: https://datascientyst.com/compare-two-excel-or-csv-files-using-pandas/ Last updated: 2025-02-10T15:47:09.000Z When working with data, you may need to compare two Excel or CSV files to: - find differences - detect updates - history changes - mistakes - validate records etc In this guide, we'll explore how to compare two Excel or CSV files using Pandas, with practical examples. ## Why Compare Files with Pandas Comparing datasets is useful when: - Validating data consistency - Detecting changes between old and new versions - Finding missing or extra rows - Comparing specific columns for modifications - Data comes from different sources Let's dive into two approaches to compare files using Pandas. ## Example 1: Compare Two CSV Files and Find Differences ### **Step 1: Load the CSV Files into Pandas** Assume we have two CSV files: - **old\_data.csv** (original dataset) - **new\_data.csv** (updated dataset) ```python import pandas as pd # Load the CSV files df_old = pd.read_csv("old_data.csv") df_new = pd.read_csv("new_data.csv") print("Old Data:\n", df_old.head()) print("\nNew Data:\n", df_new.head()) ``` We can already check if there are some differences between the rows and columnts for the files. You can find the data for both CSV files below: - `new_data.csv` ``` name,age,city Alice,25,New York Bobby,30,Los Angeles Charlie,35,Chicago David,37,San Francisco Ema,28,Houston ``` and - `old_data.csv` ``` name,age,city Alice,25,New York Bob,30,Los Angeles Charlie,35,Chicago David,40,San Francisco Emma,28,Houston ``` ### Step 2: Find Differences Between the Files To detect changes, we can use **merge** with an **outer join** and keep only the differences: ```python df_new.compare(df_old, keep_shape=False, keep_equal=True) ``` or keep the original data shape by: ```python df_new.compare(df_old, keep_shape=False, keep_equal=True) ``` result: | | name | age | | | | - | ----- | ----- | ---- | ----- | | | self | other | self | other | | 1 | Bobby | Bob | 30 | 30 | | 3 | David | David | 37 | 40 | | 4 | Ema | Emma | 28 | 28 | From the result table we can see 3 differencies: - Bobby <-> Bob - 37 <-> 40 - Ema <-> Emma If you like to find out how to highlight the changes you can check this article: [How to Compare Two Pandas DataFrames and Get Differences](https://datascientyst.com/compare-two-pandas-dataframes-get-differences/) ### Step 3: Check number of differences To check what is the number of the differences you can use: ```python ((df_old.fillna(1) == df_new.fillna(1)).melt()['value'] == 0).sum() ``` which give us 3. ## Example 2: Compare Specific Columns for Changes If both files have the same structure but contain **modified values**, we can compare specific columns. ### **Step 1: Load the Excel Files** ```python df_old = pd.read_excel("old_data.xlsx") df_new = pd.read_excel("new_data.xlsx") ``` **Note:** At this step you can compare individual sheets. For example you can download the sheet from Google Drive by: - Open the file from Google Sheets - File - Download - Comma Separated Values (.csv) ### **Step 2: Identify Modified Rows** We use `df.compare()` to highlight differences in matching rows. ```python # Compare only specific columns differences = df_old.compare(df_new) differences ``` output: | | name | age | | | | - | ----- | ----- | ---- | ----- | | | self | other | self | other | | 1 | Bobby | Bob | 30 | 30 | | 3 | David | David | 37 | 40 | | 4 | Ema | Emma | 28 | 28 | ## **Conclusion** With Pandas, comparing two CSV or Excel files is simple and effective. We explored: - **Detecting modified values using `compare()`** - **Extracting only differences** - **Counting the differences** These methods help ensure data accuracy when working with different versions of datasets. ### How to Read a Compressed CSV or JSON File in Pandas URL: https://datascientyst.com/read-compressed-csv-json-file-pandas/ Last updated: 2025-02-10T14:16:58.000Z To read compressed CSV and JSON files directly without manually decompressing them in Pandas use: **(1) Read a Compressed CSV** ```python pd.read_csv('data.csv.gz', compression='gzip') ``` **(2) Read a Compressed JSON** ```python pd.read_json('data.json.gz', compression='gzip') ``` This is useful for handling large datasets while saving storage space and improving efficiency - which can save disk space up to 10 times. ## Reading a Compressed CSV File with gzip Use the `compression` parameter in `pd.read_csv()` to read CSV files in different compressed formats. #### **Example: Reading a Gzip-Compressed CSV File** ```python import pandas as pd df = pd.read_csv('data.csv.gz', compression='gzip') df.head() ``` ## Other Supported Compression Formats You can specify different compression types as needed: - `bz2` \- Bzip2 compression - `zip` \- Zip compression - `xz` \- XZ compression ``` If the file extension matches the compression format, Pandas can automatically detect it: ```python df = pd.read_csv('data.csv.gz') # No need to specify compression ``` ## Reading a Compressed JSON File Pandas also supports reading **compressed JSON** files using `pd.read_json()`: ```python df_json = pd.read_json('data.json.gz', compression='gzip') ``` ## Multiple files found in ZIP file or nested folder If you face error like: ``` ValueError: Multiple files found in ZIP file. Only one file per ZIP: ['test/', 'test/data.csv'] ``` You will need to specify the file which has to be read: ```python import zipfile import pandas as pd with zipfile.ZipFile("/mnt/x/test.zip") as z: with z.open("test/data.csv") as f: df = pd.read_csv(f, header=0, delimiter="\t") df ``` In this case you can work with nested ZIP files or multiple files in one zip. ## Read Remote Compressed CSV file from URL link Pandas can read remote CSV files by reading the file with: `io.BytesIO(r.read())` ```python import io from urllib.request import urlopen import pandas as pd r = urlopen("https://github.com/softhints/python/raw/refs/heads/master/notebooks/csv/data.csv.zip") df = pd.read_csv(io.BytesIO(r.read()), sep=',', nrows=3) df ``` ## Detect file encoding To avoid `UnicodeDecodeError: 'ascii' codec can't decode byte 0x8b in position 1: ordinal not in range(128)` errors we can detect what is the file encoding. ```python import pandas as pd import chardet with open('/mnt/x/Datasets/data.csv', 'rb') as f: # result = chardet.detect(f.read()) # or readline if the file is not too large result = chardet.detect(f.readline()) ``` result: ``` {'encoding': 'ascii', 'confidence': 1.0, 'language': ''} ``` ### How to Save a Pandas DataFrame as a Compressed CSV/JSON File URL: https://datascientyst.com/save-pandas-dataframe-compressed-csv-json-file/ Last updated: 2025-02-10T12:40:22.000Z To save a DataFrame as a compressed CSV/JSON file using Pandas we can parameter `compression='gzip` as follows: ### CSV ```python df.to_csv('data.csv.gz', index=False, compression='gzip') ``` ### JSON ```python df.to_JSON('data.json.gz', index=False, compression='gzip') ``` ## Saving a DataFrame as a Compressed CSV You can save a Pandas DataFrame as a compressed CSV using the `compression` parameter in the `to_csv()` or `to_json()` function. Pandas supports multiple compression formats like: - gzip - bz2 - zip - xz - zstd - tar You can read more on the following link: [pandas.DataFrame.to\_csv](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.to%5Fcsv.html?ref=datascientyst.com) ## Example: Saving a DataFrame with gzip Compression Below you can do a see basic example of on-the-fly compression of the output data: ```python import pandas as pd data = {'Name': ['Alice', 'Bob', 'Charlie'], 'Age': [25, 30, 35], 'City': ['New York', 'London', 'Paris']} df = pd.DataFrame(data) df.to_csv('data.csv.gz', index=False, compression='gzip') ``` ## Other Compression Formats You can use different compression formats by changing the `compression` parameter: ```python df.to_csv('data.csv.bz2', index=False, compression='bz2') df.to_csv('data.zip', index=False, compression='zip') df.to_csv('data.csv.xz', index=False, compression='xz') ``` ## Reading the Compress CSV File To read a compressed CSV file back into a Pandas DataFrame, use `pd.read_csv()` with the `compression` parameter: ```python df = pd.read_csv('data.csv.gz', compression='gzip') ``` ## Compression Results By using compression, you can significantly reduce file size. In my tests I'm working with a file which contains the two columns: ``` https://www.example.com/south,Q6RnAzwGYA https://www.example.com/mawson,zwGYAZc https://www.example.com/sea,ZciVimr4 https://www.example.com/moo,4o6PwPjg https://www.example.com/paul,Vimr4kvJw ``` You can find the results below: - original file is 4.1 GB - Pandas compression - 689 MB - 1.5 min It took similar time for Ubuntu default compression which produced the same size - 689 MB. The advantage of Pandas is that you can exclude some columns and get smaller size after compression. ### How to Open and Convert an SQLite Database to a Pandas DataFrame URL: https://datascientyst.com/how-to-open-and-convert-an-sqlite-database-to-a-pandas-dataframe/ Last updated: 2025-01-31T22:49:27.000Z In this article, we’ll explore how to open an SQLite database and convert its tables into Pandas DataFrames with two practical examples. ## TL;DR To open and convert an SQLite database file like `data.bg` to a Pandas DataFrame we can use: ```python import sqlite3 import pandas as pd # Connect to SQLite database (or create if it doesn’t exist) conn = sqlite3.connect("example.db") # Create a cursor object cursor = conn.cursor() ``` ## Install pysqlite3 package You may need to install the [pysqlite3](https://pypi.org/project/pysqlite3/?ref=datascientyst.com) python package by: ```bash pip install pysqlite3 ``` When working with SQLite databases in Python, it’s common to extract data and analyze it using Pandas. ## Example 1: Reading an Entire Table into a DataFrame ### Step 1: Connecting to the SQLite Database First, we need to establish a connection using the `sqlite3` module in Python. ```python import sqlite3 import pandas as pd conn = sqlite3.connect("example.db") cursor = conn.cursor() # Create a cursor object ``` where file "example.db" is exported from SQlite database. ### Step 2: Creating a Sample Table (Optional) To create a sample table and insert some records. ```python cursor.execute('''CREATE TABLE IF NOT EXISTS users ( id INTEGER PRIMARY KEY, name TEXT, age INTEGER )''') cursor.executemany("INSERT INTO users (name, age) VALUES (?, ?)", [("Alice", 25), ("Bob", 30), ("Charlie", 35)]) conn.commit() ``` ### Step 3: Loading the Table into a DataFrame Now, we can load the **entire** `users` table into a Pandas DataFrame using `pd.read_sql_query()`. ```python # Read entire table into a Pandas DataFrame df = pd.read_sql_query("SELECT * FROM users", conn) df ``` the result will be: | | id | name | age | | - | -- | ------- | --- | | 0 | 1 | Alice | 25 | | 1 | 2 | Bob | 30 | | 2 | 3 | Charlie | 35 | ### Close the connection At the end we should close the sqlite3 connection afterwards with: ```python cnx.commit() cnx.close() ``` ![](https://datascientyst.com/content/images/2025/01/how-to-open-and-convert-an-sqlite-database-to-a-pandas-dataframe.png) ## Example 2: Running Custom Queries and Filtering Data Instead of loading an entire table, we can use SQL queries to retrieve only specific records. ### Step 1: Fetching Filtered Data Let’s retrieve users who are **older than 25 years**. ```python df_filtered = pd.read_sql_query("SELECT * FROM users WHERE age > 25", conn) df_filtered ``` result: | | id | name | age | | - | -- | ------- | --- | | 0 | 2 | Bob | 30 | | 1 | 3 | Charlie | 35 | ### Step 2: List all SQlite tables To list all SQlite tables with Python we can run: ```python query = """ SELECT name FROM sqlite_schema WHERE type IN ('table','view') AND name NOT LIKE 'sqlite_%' ORDER BY 1; """ pd.read_sql_query(query, conn) ``` getting only the table name: ``` `users` ``` Alternatively we can get additional table info by: ```python query = """ SELECT * FROM sqlite_master WHERE type='table' """ pd.read_sql_query(query, conn) ``` | | type | name | tbl\_name | rootpage | sql | | - | ----- | ----- | --------- | -------- | ----------------------- | | 0 | table | users | users | 2 | CREATE TABLE users ( | | | | | | | id INTEGER PRIMARY KEY, | | | | | | | name TEXT, | | | | | | | age INTEGER | | | | | | | ) | ## Conclusion Reading and converting an SQLite database to a Pandas DataFrame is a simple process using `sqlite3` and `pd.read_sql_query()`. In this guide, we explored: **Example 1:** Loading an entire table into Pandas **Example 2:** Running custom queries to fetch filtered data ### How to Use the First Row as the Header in Pandas URL: https://datascientyst.com/how-to-use-the-first-row-as-the-header-in-pandas/ Last updated: 2025-01-21T21:46:24.000Z To use the **first row as a header in Pandas** we can: **(1) Convert first row to header - reset index** ```python df.columns = df.iloc[0] df = df[1:].reset_index(drop=True) ``` **(2) Convert first row to header - keep index** ```python headers = df.iloc[0].values df.columns = headers df.drop(index=0, axis=0, inplace=True) ``` **(3) Read CSV with row 1 as header** ```python pd.read_csv('filename.csv', header = 1) ``` or: ```python pd.read_csv('filename.csv', skiprows = 1) ``` ## Steps to Use the First Row as Header ### Step 1: Create a DataFrame Assume we have a DataFrame where the first row contains the actual column headers: ```python import pandas as pd data = [ ['id', 'name', 'age'], # Row that should be the header [1, 'Alice', 25], [2, 'Bob', 30], [3, 'Charlie', 35] ] df = pd.DataFrame(data) ``` dataframe looks like: | | 0 | 1 | 2 | | - | -- | ------- | --- | | 0 | id | name | age | | 1 | 1 | Alice | 25 | | 2 | 2 | Bob | 30 | | 3 | 3 | Charlie | 35 | ### Step 2: Convert First row as Header We can use `pd.DataFrame.iloc` to extract the first row and assign it as the header: ```python df.columns = df.iloc[0] df = df[1:].reset_index(drop=True) ``` result: | | id | name | age | | - | -- | ------- | --- | | 0 | 1 | Alice | 25 | | 1 | 2 | Bob | 30 | | 2 | 3 | Charlie | 35 | ### Step 3: First row to Header and keep index We can use `pd.DataFrame.iloc` to extract the first row and assign it as the header: ```python headers = df.iloc[0].values df.columns = headers df = df.drop(index=0, axis=0) ``` result: | | id | name | age | | - | -- | ------- | --- | | 1 | 1 | Alice | 25 | | 2 | 2 | Bob | 30 | | 3 | 3 | Charlie | 35 | ### Round Pandas date to nearest year, month or week URL: https://datascientyst.com/round-pandas-date-to-nearest-year-month-or-week/ Last updated: 2024-09-04T13:02:47.000Z In this post you can find how to: - solve error: **ValueError: is a non-fixed frequency** - floor, ceil or round to year, month or week in Pandas **(1) Get year, month or week in Pandas** ```python df['year'] = df['date'].dt.year df['month'] = df['date'].dt.month df['day'] = df['date'].dt.week ``` **(2) Round or floor year in Pandas** ```python df['date'].dt.to_period('Y').dt.start_time ``` result: ``` 0 2021-01-01 1 2022-01-01 2 2023-01-01 3 2024-01-01 Name: date, dtype: datetime64[ns] ``` **(3) Round or floor year in Pandas** ```python df['date'].dt.to_period('Y').dt.end_time ``` result: ``` 0 2021-12-31 23:59:59.999999999 1 2022-12-31 23:59:59.999999999 2 2023-12-31 23:59:59.999999999 3 2024-12-31 23:59:59.999999999 Name: date, dtype: datetime64[ns] ``` **(4) Round or floor year in Pandas** ```python df['date'].dt.to_period('W-SUN').dt.start_time ``` result: ``` 0 2020-12-28 1 2022-06-27 2 2023-03-27 3 2024-10-28 Name: date, dtype: datetime64[ns] ``` **(5) Get Weekly Period in Pandas** ```python df['date'].dt.to_period('W-SUN') ``` result: ``` 0 2020-12-28/2021-01-03 1 2022-06-27/2022-07-03 2 2023-03-27/2023-04-02 3 2024-10-28/2024-11-03 Name: date, dtype: period[W-SUN] ``` **(6) Ceil or floor by offests** ```python (df['date'] + pd.offsets.YearBegin(-1)).dt.date df['date'] + pd.offsets.MonthBegin(-1) ``` result: ``` 0 2020-01-01 1 2022-01-01 2 2023-01-01 3 2024-01-01 Name: date, dtype: object ``` ## ValueError: is a non-fixed frequency There is open issue about using dt.ceil, dt.floor, dt.round with coarser target frequencies: - YE - ME - W [Can't use dt.ceil, dt.floor, dt.round with coarser target frequencies](https://github.com/pandas-dev/pandas/issues/15303?ref=datascientyst.com) Usually you will receive error like: - `ValueError: <30 * YearEnds: month=12> is a non-fixed frequency` - `ValueError: is a non-fixed frequency` - `ValueError: is a non-fixed frequency` - `ValueError: is a non-fixed frequency` To solve this error you can: - get the period by: `df['date'].dt.to_period('W-SUN')` - get the start or the end of the period: `.dt.start_time` or `.dt.end_time` ## Sample Data ```python import pandas as pd data = { 'date': [ pd.Timestamp('2021-01-01 11:04:45'), pd.Timestamp('2022-07-01 12:035:15'), pd.Timestamp('2023-04-01 12:16:30'), pd.Timestamp('2024-11-01 15:57:59') ] } df = pd.DataFrame(data) df ``` ### Rounding Pandas Timestamps to the Nearest Minute, 15M, 30M URL: https://datascientyst.com/rounding-pandas-timestamps-to-the-nearest-minute-15m-30m/ Last updated: 2024-09-04T12:38:05.000Z Below you can find multiple ways to round to the nearest minute or seconds in Pandas: **(1) rounding pandas timestamp** ```python pd.to_datetime(df['date']).dt.round(freq='30min').dt.time ``` result: ``` 0 11:00:00 1 12:30:00 2 12:00:00 3 15:30:00 Name: date, dtype: object ``` **(2) round date to 30 minutes python** ```python df['date'].dt.round('15min') ``` results: ``` 0 2024-09-01 11:00:00 1 2024-09-01 12:30:00 2 2024-09-01 12:15:00 3 2024-09-01 15:45:00 Name: date, dtype: datetime64[ns] ``` **(3) rounding pandas timestamps to the nearest minute or 5 minutes** ```python df['date'].dt.round('5min') ``` ``` 0 2024-09-01 11:05:00 1 2024-09-01 12:35:00 2 2024-09-01 12:15:00 3 2024-09-01 16:00:00 Name: date, dtype: datetime64[ns] ``` **(4) round to a minute in python/pandas** ```python df['date'].dt.round("10s") ``` result: ``` 0 2024-09-01 11:04:40 1 2024-09-01 12:35:20 2 2024-09-01 12:16:30 3 2024-09-01 15:58:00 Name: date, dtype: datetime64[ns] ``` **(4) floor to a closest minute** ```python df['date'].dt.floor('30min') ``` result: ``` 0 2024-09-01 11:00:00 1 2024-09-01 12:30:00 2 2024-09-01 12:00:00 3 2024-09-01 15:30:00 Name: date, dtype: datetime64[ns] ``` **(5) ceil to a closest minute** ```python df['date'].dt.ceil('30min') ``` result: ``` 0 2024-09-01 11:30:00 1 2024-09-01 13:00:00 2 2024-09-01 12:30:00 3 2024-09-01 16:00:00 Name: date, dtype: datetime64[ns] ``` ## Different frequencies to round - `d` \- day - `h` \- hour - `min` \- minute - `s` \- second ## Sample Data ```python import pandas as pd data = { 'date': [ pd.Timestamp('2024-09-01 11:04:45'), pd.Timestamp('2024-09-01 12:035:15'), pd.Timestamp('2024-09-01 12:16:30'), pd.Timestamp('2024-09-01 15:57:59') ] } df = pd.DataFrame(data) df ``` ## Resource - [Datetimelike rounding](https://pandas.pydata.org/docs/whatsnew/v0.18.0.html?ref=datascientyst.com#datetimelike-rounding) - [pandas.Timestamp.round](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Timestamp.round.html?ref=datascientyst.com) ### Guide: How To Move a Column to the Front in Pandas DataFrame? URL: https://datascientyst.com/guide-move-column-front-pandas-dataframe/ Last updated: 2024-04-20T10:12:10.000Z In this guide, you can learn how to **move columns by name to the front in Pandas DataFrame**. There are situations where you might need to move a specific column to the front of the DataFrame: - for better visibility - to facilitate further analysis - to fulfill business requirements - export purposes Let's guide you through the process of moving a column by name to the front of a Pandas DataFrame, helping you organize your data more effectively. For sorting column in DataFrame check: [How to Change the Order of Columns in Pandas DataFrame](https://datascientyst.com/change-order-columns-pandas-dataframe/) ![](https://datascientyst.com/content/images/2024/04/move-column-front-pandas-dataframe.webp) ## Data To illustrate the example let's create simple Pandas DataFrame: ```python import pandas as pd # Sample DataFrame data = { 'id': [1, 2, 3, 4, 5], 'name': ['John', 'Alice', 'Bob', 'Charlie', 'Emma'], 'value': [10, 20, 30, 40, 50], 'category': ['A', 'B', 'C', 'B', 'A'] } df = pd.DataFrame(data) df ``` | | id | name | value | category | | - | -- | ------- | ----- | -------- | | 0 | 1 | John | 10 | A | | 1 | 2 | Alice | 20 | B | | 2 | 3 | Bob | 30 | C | | 3 | 4 | Charlie | 40 | B | | 4 | 5 | Emma | 50 | A | ## Move Column by `pop()` and `insert()` To move a specific column to the front of the DataFrame, you can use: Pandas `pop()` and `insert()` functions: ```python col = 'category' column_to_move = df.pop(col) df.insert(0, col, column_to_move) ``` Column category will be moved as first column in the DataFrame: | | category | id | name | value | | - | -------- | -- | ------- | ----- | | 0 | A | 1 | John | 10 | | 1 | B | 2 | Alice | 20 | | 2 | C | 3 | Bob | 30 | | 3 | B | 4 | Charlie | 40 | | 4 | A | 5 | Emma | 50 | ## Move Column by selecting a list of columns As an alternative solution we can use Python list comprehension to predefine the column order. To move single column to front of the DataFrame by using list comprehension we can: ```python col_to_move = 'category' df[[col_to_move] + [ col for col in df.columns if col != col_to_move ]] ``` The result is the same as the on using `pop` ## Move Multiple Columns To move multiple Pandas columns to the start of the DataFrame we can build a list with the desired column order: ```python cols_to_move = ['category', 'name'] df[ cols_to_move + [ col for col in df.columns if col not in cols_to_move ]] ``` Now the columns `'category', 'name'` are move at the beginning: | | category | name | id | value | | - | -------- | ------- | -- | ----- | | 0 | A | John | 1 | 10 | | 1 | B | Alice | 2 | 20 | | 2 | C | Bob | 3 | 30 | | 3 | B | Charlie | 4 | 40 | | 4 | A | Emma | 5 | 50 | ## Conclusion Reordering columns in a Pandas DataFrame allows you to organize your data effectively for analysis, sharing and visualization. Hopefully, now you can easily move a **specific columns to the front of the DataFrame.** ## Resources: Additional resources on reordering and sorting columns in Pandas: - [How to Change the Order of Columns in Pandas DataFrame](https://datascientyst.com/change-order-columns-pandas-dataframe/) - [Move column by name to front of table in pandas](https://stackoverflow.com/questions/25122099/move-column-by-name-to-front-of-table-in-pandas?ref=datascientyst.com) ### Guide: How to Exclude Rows while Sorting a DataFrame in Pandas URL: https://datascientyst.com/exclude-rows-while-sorting-dataframe-pandas/ Last updated: 2024-04-20T08:59:54.000Z In this short how to guide you can learn how to **exclude rows from sorting in Pandas.** To keep the original order of specific rows is useful when you deal with: \* multi-dimensional data - filtering out outliers - excluding specific observations from the sorting process. Let's explore how to sort a DataFrame in Pandas while omitting specific rows. ![](https://datascientyst.com/content/images/2024/04/2.-Sort-and-exclude.webp) ## Step 1: Data First, let's create a sample dataframe which will help us to illustrate the example better ```python import pandas as pd # Sample DataFrame data = { 'id': ['ean', '-', 1, 2, 3, 4, 5], 'name': ['full', '-','John', 'Alice', 'Bob', 'Charlie', 'Emma'], 'value': ['$', '-',10, 20, 30, 40, 50], 'category': ['A-C', '-','A', 'B', 'C', 'B', 'A'] } df = pd.DataFrame(data) df ``` Output is: | | id | name | value | category | | - | --- | ------- | ----- | -------- | | 0 | ean | full | $ | A-C | | 1 | \- | \- | \- | \- | | 2 | 1 | John | 10 | A | | 3 | 2 | Alice | 20 | B | | 4 | 3 | Bob | 30 | C | | 5 | 4 | Charlie | 40 | B | | 6 | 5 | Emma | 50 | A | ## Step 2: Define exclude conditions First we will define conditions based on which we will decide which rows to be excluded from the sorting by: ```python condition = (df['id'] == 'ean') | (df['id'] == '-') excluded = df[condition] included = df[~condition] ``` ## Step 3: Sorting Next use the `sort_values()` function to sort the included DataFrame according to your desired criteria. Suppose you want to sort by the "name" column in ascending order: ```python sorted = included.sort_values(by="name",ascending=True) ``` ## Step 4: Merge results Finally we can merge the unsorted and sorted data by: ```python pd.concat([excluded, sorted]) ``` ## Full Example Code: Sort and Exclude Full example showing how to exclude specific rows from the sorting process, you can filter them out before sorting. ```python condition = (df['id'] == 'ean') | (df['id'] == '-') excluded = df[condition] included = df[~condition] sorted = included.sort_values(by="name",ascending=True) df = pd.concat([excluded, sorted]) ``` the final output is keep the first rows without a change while sort the rest by `name`: | | id | name | value | category | | - | --- | ------- | ----- | -------- | | 0 | ean | full | $ | A-C | | 1 | \- | \- | \- | \- | | 3 | 2 | Alice | 20 | B | | 4 | 3 | Bob | 30 | C | | 5 | 4 | Charlie | 40 | B | | 6 | 5 | Emma | 50 | A | | 2 | 1 | John | 10 | A | This example shows how to exclude the first N rows. You can tailor it for excluding the last N rows by changing the concat order. ## Conclusion In this guide we saw how to **sort DataFrames in Pandas while excluding specific rows**. By following the steps outlined above, you can ensure your data is sorted accurately according to your criteria, while omitting any rows that may not be relevant to your analysis. ## Resources - [Pandas Python: sort dataframe but don't include given row](https://stackoverflow.com/questions/26220681/pandas-python-sort-dataframe-but-dont-include-given-row?ref=datascientyst.com) ### How to Merge with Missing Values in Pandas URL: https://datascientyst.com/how-to-merge-with-missing-values-in-pandas/ Last updated: 2024-04-13T08:56:07.000Z In this article, you can learn **how to merge DataFrames in Pandas with handling of missing values.** When working with data in Pandas, you often need to merge multiple DataFrames to consolidate information from different sources. However, not all data align perfectly, and missing values are common. Missing values can cause unexpected results or performance issues. ![How to Merge with Missing Values in Pandas](https://datascientyst.com/content/images/2024/04/how-to-merge-with-missing-values-in-pandas.webp) ## Create sample data First let's create two sample DataFrames which can be used to illustrate the merging example: ```python import pandas as pd import numpy as np foo = pd.DataFrame([ ['a',1,2], ['b',4,5], ['c',7,8], [np.NaN,10,11] ], columns=['id','x','y']) display(foo) bar = pd.DataFrame([ ['a',3], ['c',9], [np.NaN,12] ], columns=['id','z']) display(bar) ``` data for foo: | | id | x | y | | - | --- | -- | -- | | 0 | a | 1 | 2 | | 1 | b | 4 | 5 | | 2 | c | 7 | 8 | | 3 | NaN | 10 | 11 | and bar: | | id | z | | - | --- | -- | | 0 | a | 3 | | 1 | c | 9 | | 2 | NaN | 12 | ## merge default bahavior By default pandas will merge the the missing values from the DataFrames. ```python pd.merge(foo, bar, how='inner', on='id') ``` So merging by column `id` will give us 3 rows instead of 2 ( which differs from the most DB): | | id | x | y | z | | - | --- | -- | -- | -- | | 0 | a | 1 | 2 | 3 | | 1 | c | 7 | 8 | 9 | | 2 | NaN | 10 | 11 | 12 | ## merge without NAN / missing values To merge without the missing values from the merge column we can drop the values from the first dataframe: ```python pd.merge(foo.dropna(subset=['id']), bar, how='inner', on='id') ``` This will give us: | | id | x | y | z | | - | -- | - | - | - | | 0 | a | 1 | 2 | 3 | | 1 | c | 7 | 8 | 9 | ## merge with missing values To keep all values after merging we can use `how='outer'`: ```python pd.merge(foo.dropna(subset=['id']), bar, how='outer', on='id') ``` Finally we have data from both dataframes including the missing values: | | id | x | y | z | | - | --- | -- | -- | ---- | | 0 | a | 1 | 2 | 3.0 | | 1 | b | 4 | 5 | NaN | | 2 | c | 7 | 8 | 9.0 | | 3 | NaN | 10 | 11 | 12.0 | ## Conclusion **Merging DataFrames with proper handling of missing values in Pandas** is a fundamental operation in data analysis. By using the `merge()` function with different join parameters and leveraging Pandas' capabilities to handle missing values, you can efficiently consolidate data from multiple sources. ### How to Read CSV or JSON from URL With Authentication in Pandas URL: https://datascientyst.com/how-to-read-csv-from-url-with-authentication-in-pandas/ Last updated: 2024-04-13T09:43:11.000Z Reading a **CSV file directly from a URL into Pandas** is a common task, especially when dealing with web data. However, sometimes the data you need requires authentication to access. Fortunately, Python and Pandas provide straightforward methods to handle this scenario. Let's explore how to read a CSV from a URL with authentication using Pandas. ## Authentication with requests If the URL requires authentication, you'll need to provide credentials to access it. This typically involves passing a username and password or an access token. Python requests supports various authentication methods, including HTTP basic authentication and token-based authentication. For example, if the URL requires basic authentication, you can use the requests library to pass the credentials: ```python import requests import pandas as pd from io import StringIO from requests.auth import HTTPBasicAuth url = 'https://example.com/items/data.json' user = 'username' password = 'password' data = requests.get(url, auth=HTTPBasicAuth(user, password)) df = pd.read_json(StringIO(data.text), lines=True) df ``` the same applies for `read_csv` method. ## Authentication with custom headers Pandas offers parameter `storage_options` for: - `read_csv` - `read_json methods` We can use custom headers to provide authentication information to pandas by generating authentication header like: `{'Authorization': 'Basic xxxx'}` ### read\_json + storage\_options The following example provides how this can be done: ```python from http.client import HTTPSConnection from base64 import b64encode import base64 def basic_auth(username, password): token = b64encode(f"{username}:{password}".encode('utf-8')).decode("ascii") return f'Basic {token}' username = "username" password = "password" headers = { 'Authorization' : basic_auth(username, password) } headers df = pd.read_json( "https://example.com/items/data.json", storage_options=headers, lines=True ) ``` ### read remote CSV file Below you can find shorter version for `read_csv` method: ```python import pandas as pd from base64 import b64encode df = pd.read_csv( 'https://example.com/items/data.csv', storage_options={'Authorization': b'Basic %s' % b64encode(b'username:password')}) df ``` ## Read files from amazon buckets ### public bucket To read data from amazon public buckets we need first to install library by - `s3fs`: ```python pip install s3fs ``` then we can use: ```python import pandas as pd pd.read_csv( "s3://ncei-wcsd-archive/data/processed/SH1305/18kHz/SaKe2013" "-D20130523-T080854_to_SaKe2013-D20130523-T085643.csv", storage_options={"anon": True} ) ``` ### private buckets as alternative we can use Pandas `read_csv` method with AWS accounts to read remote data: ```python df = pd.read_csv("s3://my-private-bucket/data.csv") ``` We can pass the amazon keys and secrets by: ```python df = pd.read_csv( "s3://my-private-bucket/data.csv", storage_options={"key": "AKIAIOSFODNN7EXAMPLE", "secret": "SECRET"}, ) ``` or by using [boto3](https://pypi.org/project/boto3/?ref=datascientyst.com) library: ```python import os import pandas as pd import boto3 session = boto3.Session(profile_name="test") os.environ['AWS_ACCESS_KEY_ID'] = session.get_credentials().access_key os.environ['AWS_SECRET_ACCESS_KEY'] = session.get_credentials().secret_key df = pd.read_csv("s3://xxxx.csv") ``` ## Conclusion Reading a CSV from a URL with authentication in Pandas is a straightforward process. By following the steps outlined above, you can access data hosted on the web securely and leverage the powerful data manipulation capabilities of Pandas for your analysis. ## Resources You can learn more about reading remote files with Pandas here: - [read\_csv](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) - [read\_json](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read%5Fjson.html?ref=datascientyst.com) - [Reading/writing remote files](https://pandas.pydata.org/docs/user%5Fguide/io.html?ref=datascientyst.com#reading-writing-remote-files) ### ImportError: matplotlib is required for plotting when the default backend "matplotlib" is selected URL: https://datascientyst.com/pandas-importerror-matplotlib-is-required-for-plotting/ Last updated: 2024-01-20T23:24:52.000Z Pandas, a powerful data manipulation library in Python, offers a convenient way to analyze and visualize data through its integration with the Matplotlib plotting library. However, users may encounter an ImportError when attempting to use Pandas for plotting, specifically indicating that Matplotlib is required. In this article, we will explore the issue and provide solutions to address the Pandas ImportError related to Matplotlib. ## Understanding the ImportError: matplotlib is required for plotting The error message typically looks like this: ``` ImportError: matplotlib is required for plotting when the default backend "matplotlib" is selected ``` This error arises when Pandas attempts to use Matplotlib for plotting, but Matplotlib is either not installed or is not accessible. ```python import pandas as pd ser = pd.Series([1, 2, 3, 3]) plot = ser.plot(kind='hist', title="My plot") ``` ## Solution: Install Matplotlib The most straightforward solution is to ensure that Matplotlib is installed. Use the following command to install Matplotlib via pip: ```python pip install matplotlib ``` This command installs the Matplotlib library along with its dependencies, allowing Pandas to use it for plotting functionalities. The package information is available on: [https://pypi.org/project/matplotlib/](https://pypi.org/project/matplotlib/?ref=datascientyst.com) P.S. Note that matplotlib is part of the optional dependencies for Pandas. You can find all of them here: [Pandas Visualization Dependencies](https://pandas.pydata.org/docs/getting%5Fstarted/install.html?ref=datascientyst.com#visualization) ## Check Matplotlib Version It's essential to have a compatible version of Matplotlib installed. Pandas may have specific version requirements. To install a specific version of Matplotlib, use: ```bash pip install matplotlib==3.8.0 ``` Replace `3.8.0` with the version you want to install. ## Import Matplotlib in Your Script In some cases, the error may persist if Matplotlib is not imported explicitly in your script. Ensure that you have the following import statement at the beginning of your Python script or Jupyter Notebook - `import matplotlib.pyplot as plt`: ```python import matplotlib.pyplot as plt import pandas as pd ser = pd.Series([1, 2, 3, 3]) plot = ser.plot(kind='hist', title="My plot") ``` This statement ensures that Matplotlib is accessible for Pandas when performing plotting operations. ## Check Virtual Environment If you are working within a virtual environment, make sure it is activated before installing Matplotlib. If the virtual environment is not active, the installation might be done in the global Python environment instead. ```python pip freeze ``` will list the installed python packages. ## Upgrade Pandas Ensure that you have the latest version of Pandas installed. Use the following command to upgrade Pandas: ```python pip install --upgrade pandas ``` Upgrading to the latest version may resolve compatibility issues with Matplotlib. ## Conclusion The Pandas Error: **"ImportError: matplotlib is required for plotting when the default backend "matplotlib"** is selected" is commonly encountered when Matplotlib is not properly installed or is incompatible with the version of Pandas being used. By following the steps outlined above, you can resolve this error and unlock Pandas plotting capabilities with Matplotlib. ### Importerror missing optional dependency html5lib pandas example URL: https://datascientyst.com/importerror-missing-optional-dependency-html5lib-pandas-example/ Last updated: 2024-01-20T21:33:10.000Z In this tutorial, we'll show how to solve a common **Pandas error – "importerror missing optional dependency html5lib pandas example"**. We get this error from the Pandas when we try to use the method `read_html()` but the library `html5lib` is not installed. ## Fix Importerror missing optional dependency html5lib This library can be installed by: ```python pip install html5lib ``` ## Dependency html5lib The html5lib library is optional dependency for Pandas. The full list of optional dependencies for Pandas can be found here: [Pandas Optional dependencies](https://pandas.pydata.org/pandas-docs/stable/getting%5Fstarted/install.html?ref=datascientyst.com#optional-dependencies) You can find more information for this library on link: - [https://github.com/html5lib/html5lib-python](https://github.com/html5lib/html5lib-python?ref=datascientyst.com) - [https://pypi.org/project/html5lib/](https://pypi.org/project/html5lib/?ref=datascientyst.com) ## read\_html dependencies There are 3 dependencies recommended for Pandas when you need to read data with method `read_html`: One of the following combinations of libraries is needed to use the top-level [read\_html()](../reference/api/pandas.read%5Fhtml.html#pandas.read%5Fhtml "pandas.read_html") function: - [BeautifulSoup4](https://www.crummy.com/software/BeautifulSoup?ref=datascientyst.com) and [html5lib](https://github.com/html5lib/html5lib-python?ref=datascientyst.com) - [BeautifulSoup4](https://www.crummy.com/software/BeautifulSoup?ref=datascientyst.com) and [lxml](https://lxml.de/?ref=datascientyst.com) - [BeautifulSoup4](https://www.crummy.com/software/BeautifulSoup?ref=datascientyst.com) and [html5lib](https://github.com/html5lib/html5lib-python?ref=datascientyst.com) and [lxml](https://lxml.de/?ref=datascientyst.com) They can be installed by: ```python pip install "pandas[html] ``` More info: [Pandas HTML](https://pandas.pydata.org/pandas-docs/stable/getting%5Fstarted/install.html?ref=datascientyst.com#html) ## Older Pandas Version For older Pandas and Python version the same error: > importerror missing optional dependency html5lib pandas example Can be shown even if the library is installed. The reason might be wrong path to the local file: ```python import pandas as pd df = pd.read_html("/test/test.htm") ``` Latest Pandas version gives error: `ValueError: No tables found` when the file is missing or there's no table ### AttributeError: 'DataFrame' object has no attribute 'append' - Pandas URL: https://datascientyst.com/fix-attributeerror-dataframe-object-has-no-attribute-append-pandas/ Last updated: 2024-01-19T23:04:36.000Z In this tutorial, we'll see how to solve a Pandas error: **AttributeError: 'DataFrame' object has no attribute 'append'**. We will also answer on the questions: - Why is append not working in pandas? - How do I fix pandas attribute error? - How do you append an object to a DataFrame in Python? - How to append a new row to Pandas DataFrame? ## Why is append not working in pandas? Pandas method `append()` was deprecated in version 1.4. So in more recent versions of Pandas 2.0, `append()` is no longer used to add rows to a DataFrame. You can achieve the same result using the `concat()` method. For more information you can find: [pandas.DataFrame.append](https://pandas.pydata.org/pandas-docs/version/1.4/reference/api/pandas.DataFrame.append.html?ref=datascientyst.com) ![](https://datascientyst.com/content/images/2024/01/fix-attributeerror-dataframe-object-has-no-attribute-append-pandas.opti.webp) ## AttributeError: 'DataFrame' object has no attribute 'append' The error: AttributeError: 'DataFrame' object has no attribute 'append' can be reproduced by the following example. We would like to append new data to DataFrame: | | A | B | | - | - | - | | 0 | 1 | 2 | | 1 | 3 | 4 | The code is: ```python import pandas as pd data = {'A': [1, 2], 'B': [3, 4]} df = pd.DataFrame(data) df = df.append({'A': 5, 'B': 6}) ``` this results into error: > AttributeError: 'DataFrame' object has no attribute 'append' ## Fix for AttributeError - concat To fix the error we can use the method `concat()` to append two DataFrames in Pandas. So the append syntax: ```python df.append({'A': 5, 'B': 6}) ``` should be changed to: ```python pd.concat([df, pd.DataFrame([{'A': 5, 'B': 6}])]) ``` or using parameter - `ignore_index=True`: ```python pd.concat([df, pd.DataFrame([{'A': 5, 'B': 6}])], ignore_index=True) ``` The above example will become: ```python import pandas as pd data = {'A': [1, 3], 'B': [2, 4]} df = pd.DataFrame(data) pd.concat([df, pd.DataFrame([{'A': 5, 'B': 6}])], ignore_index=True) ``` Final result is appended new row to the original DataFrame: | | A | B | | - | - | - | | 0 | 1 | 2 | | 1 | 3 | 4 | | 2 | 5 | 6 | ## Alternative Solutions In this section you can find several alternative solutions: - downgrade to older Pandas version - `df = df1._append(df2,ignore_index=True)` - `df.loc[len(df)] = new_row # only use with a RangeIndex!` - `pd.concat([df, df_extended])` more information and details on that error can be found on this link: [Error "'DataFrame' object has no attribute 'append'"](https://stackoverflow.com/questions/75956209/error-dataframe-object-has-no-attribute-append?ref=datascientyst.com) ## Conclusion To sum up, this article shows how using method concat can solve the "AttributeError: 'DataFrame' object has no attribute 'append'" Python error. ### How to Keep the First Value of Column After Explode in Pandas? URL: https://datascientyst.com/keep-first-value-column-after-explode-pandas/ Last updated: 2024-01-17T22:38:05.000Z In this quick tutorial, we're going to look at how to keep the first value of column after explode in Pandas? Suppose we have a DataFrame which has a column with nested data - list or JSON. Let's work with the following DataFrame: ```python import pandas as pd data = {'ID': [1, 2, 3], 'Items': [['A', 'B'], ['C', 'D'], ['E', 'F', 'G']]} df = pd.DataFrame(data) ``` Data looks like: | | ID | Items | | - | -- | ----------- | | 0 | 1 | \[A, B\] | | 1 | 2 | \[C, D\] | | 2 | 3 | \[E, F, G\] | Explode the column items will return mulitple items per row: ```python ``` result: | | ID | Items | | - | -- | ----- | | 0 | 1 | A | | 0 | 1 | B | | 1 | 2 | C | | 1 | 2 | D | | 2 | 3 | E | | 2 | 3 | F | | 2 | 3 | G | Notice that index contains duplicates for each item present in the original column ## Explode List Column Keep First Item To explode column Items and keep only the first item per each row we can drop duplicates: ```python d = df.explode('Items') d[~d.index.duplicated()] ``` result: | | ID | Items | | - | -- | ----- | | 0 | 1 | A | | 1 | 2 | C | | 2 | 3 | E | Alternative simpple solution for List columns is the following: ```python df['Items'].apply(lambda x: x[0] if isinstance(x, list) else x) ``` The result is the same ## Explode JSON column - keep first item For JSON columns we can use the following code. Let say that we work with column `'tags'` which contains JSON data like: `[{'id': '62b5ac97cdb6600403c69f0f', 'name': 'Cheat Sheet', 'slug': '108-cheat-sheet', 'created_at': '2022-06-24T12:22:47.000Z'...}]` This column can be exploded and set to the original DataFrame by: ```python dd = df.explode('tags')['tags'] dd = pd.json_normalize(dd[~dd.index.duplicated()]) df[['tag', 'tag_published_at']] = dd[['slug', 'created_at']] ``` This will extract the complex structure and get only specfic columns and values. If you need to explode given column and set default value you can use mask: ```python import pandas as pd data = {'ID': [1, 2, 3], 'Items': [['A', 'B'], ['C', 'D'], ['E', 'F', 'G']], 'val': [5, 10, 7]} df = pd.DataFrame(data) ``` data: | | ID | Items | val | | - | -- | ----------- | --- | | 0 | 1 | \[A, B\] | 5 | | 1 | 2 | \[C, D\] | 10 | | 2 | 3 | \[E, F, G\] | 7 | ## Explode and Mask Expanding with setting value for another column we can do: ```python d = df.explode('Items') d['val'] = d['val'].mask(d['ID'].duplicated(), 0) d ``` results in: | | ID | Items | val | | - | -- | ----- | --- | | 0 | 1 | A | 5 | | 0 | 1 | B | 0 | | 1 | 2 | C | 10 | | 1 | 2 | D | 0 | | 2 | 3 | E | 7 | | 2 | 3 | F | 0 | | 2 | 3 | G | 0 | Now we can keep first or different item by: ```python d[d['val'] != 0] ``` ## Conclusion By using a lambda function or Pandas functions, you can preserve the first value of the specified column even after exploding it. This ensure that your data maintains its original structure while benefitting from the exploded format. ### Pandas vs R - cheat sheet URL: https://datascientyst.com/pandas-vs-r-cheat-sheet/ Last updated: 2023-12-02T22:34:21.000Z This is a **Python/Pandas vs R cheatsheet** for a quick reference for switching between both. The post contains **equivalent operations between Pandas and R**. The post includes the most used operations needed on a daily baisis for data analysis. Have in mind that some examples might differ due to different indexing or updates. If you want to contribute feel free to suggest changes or additions on GitHub: [pandas\_r\_cheatsheet.csv](https://github.com/softhints/Pandas-Tutorials/blob/master/cheatsheet/pandas%5Fr%5Fcheatsheet.csv?ref=datascientyst.com) ## Pandas vs R cheatsheet Single column 2 columns 3 columns Hide navigation Hide TOC ## Setup Import and package installation ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_basics.png) `import pandas as pd import numpy as np` library(dplyr) library(ggplot2) # error if missing require(ggplot2) # warning Import libraries and modules `pip install pandas` install.packages('ggplot2') install package `https://pypi.org/` https://cran.r-project.org/web/packages/ Search Packages ## Data Structures Pandas Series vs R Array DataFrame comparison ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_data_structures-1.png) `s = pd.Series(np.arange(5))` s <- 0:4 # > Pandas series vs R vectors `s[0]` s\[0\] Get first element of array or Series `df = pd.DataFrame( {'col_1': [11, 12, 13], 'col_2': [21, 22, 23]}, index=[0, 1, 3])` df = data.frame ( col\_1 = c(11, 12, 13), col\_2 = c(21, 22, 23) ) rownames(df) <- c(0,1,3) # > Pandas vs R DataFrame `import numpy as np import pandas as pd data = np.random.randn(10, 3) cols = list('abc') pd.DataFrame(data, columns=cols)` data.frame(a=rnorm(10), b=rnorm(10), c=rnorm(10)) Create random DataFrame ## Read Import Data R vs Pandas ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_read.png) `df = pd.read_csv('file.csv')` df <- read.csv('file.csv') # > Read CSV file `pd.read_json('file.json')` library(jsonlite) df <- read\_json('file.json') # > Read JSON file `pd.read_csv('https://example.com/file.csv')` read.csv(url('https://example.com/file.csv')) Read data from URL `df = pd.read_fwf('delim_file.txt')` df <- read\_fwf('delim\_file.txt') # > Read delimited file ## Write Data export - Pandas vs R ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_write.png) `df.to_csv('file.csv')` write.csv(df, 'data.csv', row.names=FALSE) Writes to a CSV file `df.to_json('file.json')` js\_file <- jsonlite::toJSON(df2, pretty = TRUE) write(js\_file, 'file.json') # > Writes to a file in JSON format ## Inspect Data Statistics, samples and summary of the data ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_info.png) `df.shape` dim(df) return dimensions `df.head(6)` head(df, 6) First n rows `df.tail(6)` tail(df, 6) Last n rows `df.describe()` summary(df) Summary statistics `df.loc[:, :'a'].describe()` summary(df\[, 'a'\]) Describe columns `df['A'].mean()` mean(df\[, 'a'\]) Statistical functions ` df.sample(n=10)` sample\_n(df, 10) Sample n random rows ## Select Select data by index, by label, get subset ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_select.png) `df.loc[1:3, :]` df\[2:4,\] Select first N rows - all columns `df.loc[[1, 2, 3], :]` df\[c(2,3,4),\] Select rows by index `df.loc[:, ['a', 'b']].copy()` copy <-data.frame(df\[,c('a','b')\]) # > Select columns by name(copy) `df.loc[:, ['a']]` df\[, 'a'\] Select columns by name(reference) `df.loc[1:3, ['b', 'a']]` df\[2:4, c('b','a')\] Subset rows and columns `df.loc[[3,1], ['b', 'a']]` df\[4:2, c('b','a')\] Reverse selection `df[df['a'].isna()]` df\[is.na(df$a), \] Select NaN values `df['a'].dropna()` df\[!is.na(df$a), \] Select non NaN values ## Add rows/columns Add new columns and rows ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_add.png) `df['new col'] = df['col'] * 100` df$new <- df\[, 'a'\] \* 100 # > Add new column based on other column `df['new col'] = False` df$new <-FALSE # > Add new column single value `df.loc[-1] = [1, 2, 3]` df\[nrow(df) + 1,\] = c(1,2,3) Add new row at the end of DataFrame `df.append(df2, ignore_index = True)` rbind(df, df2) add rows from DataFrame to existing DataFrame ## Drop rows/columns/nan Drop data from DataFrame ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_drop.png) `s.drop(1)` s\[!(s == 1)\] (Series) Drop values from Series by index (row axis) `s.drop([1, 2])` s\[!(s %in% c(1,2))\] (Series) Drop values from Series by index (row axis) `df.drop('b' , axis=1) ` subset(df, select = -c(b)) Drop column by name col\_1 (column axis) `df.dropna()` library(tidyr) df %>% drop\_na() Drops all rows that contain null values `df.dropna(axis=1)` janitor::remove\_empty(df, which = 'cols') Drops all columns that contain null values ## Sort values/index Sorting and rank values in Pandas vs R ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_sort.png) `sorted([2,3,1])` sort(c(2,3,1)) sort array of values `sorted([2,3,1], reverse=True)` sort(c(2,3,1),decreasing=TRUE) sort in reverse order `df['a'].sort_values()` sort(df\[, 'a'\]) sort DataFrame by column `df.sort_values(['a', 'b'], ascending=[False, True])` df\[order(-df$a, df$b), \] sort DataFrame by multiple columns ## Filter Filter data based on multiple criteria ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_filter.png) `df.loc[:, df.isna().any()]` apply(df, 2, function(x) any(is.na(x))) find columns with na `df.loc[df.isna().any(), :]` apply(df, 1, function(x) any(is.na(x))) find rows with na `df[df['col_1'] > 100]` filter(df, col\_1 > 100) Values greater than X `df[(df['a']=='a')&(df['b']>=10)]` filter(df, a == 'a', b > 10) Filter Multiple Conditions - & - and; | - or `df[df['a'] == 'test']` filter(df, a == 'test') filter by sting value `df[(df['a'] == 'test') & (df['b'] == 'a2') ]` filter(df, a == 'test', b == 'a2' ) combine conditions ## Group by Group by and summarize data ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_groupby.png) `df.groupby('a')` group\_by(df, 'a') Group by single column `df.groupby(['a', 'b']).c.sum()` aggregate(df$b, by=list(a=df$a), FUN=sum) group by multiple columns and sum third `df['a'].value_counts()` dplyr::count(df, a, sort = TRUE) group by and count ## Convert Convert to date, string, numeric ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_convert.png) `df['a'].fillna(0)` library(dplyr) df <- df %>% mutate(a = if\_else(is.na(a), 0, a)) # > replace NA values `df.replace('..', None)` df\[df == '..'\] <- NA # > convert .. to NA `df['col_1'].astype('int64')` strtoi(c('1', '2'), base = 0L) convert string to int `pd.to_datetime(df['date'], format='%Y-%m-%d')` dates <- c('2023-09-04', '2023-09-06') as.Date(dates, format='%Y-%m-%d') # > convert string to date P.S. Due to bug in the blog platform `<-` is displayed with R comment. So instead of: `s <- 0:4` the code is shown as `s <- 0:4 #>` ## 0\. How to Install R Packages To install new packages in R follow these steps: - Launch your R console or RStudio. - Install single package - `install.packages('jsonlite')` - To install multiple packages simultaneously: - `install.packages(c('jsonlite', 'ggplot2'))` - R will download and install the specified packages from the CRAN (Comprehensive R Archive Network) repository. Once the installation is complete, you can load the package into your R session using the `library('jsonlite')` function. ### Install ggplot2 in R For example, to install the "ggplot2" package, you can use the commands: ```python install.packages('jsonlite') library('jsonlite') ``` ## 1\. Main Differences: R and Pandas Pandas and R are both popular tools/languages for data analysis, manipulation and statistics. Some key differences between them: ### Indexing One big difference between R and Pandas is indexing: - R - 1 based \* [Indexing from zero in R](https://www.r-bloggers.com/2021/12/indexing-from-zero-in-r/?ref=datascientyst.com) \* [Package ‘index0’](https://cran.r-project.org/web/packages/index0/index0.pdf?ref=datascientyst.com) \* Pandas - 0 based ### Syntax - R syntax is tailored for statistical analysis. It uses functions and operators that are well-suited for data manipulation, statistics and visualization. - Pandas uses Python syntax, which is more general-purpose. It leverages Python's data structures like DataFrames and Series for data manipulation. Pandas also use the indexing, slicing and other Python techniques. Below you can compare the creation of DataFrames in Pandas vs R: ```python # pandas import pandas as pd df = pd.DataFrame(np.random.randn(10, 5), columns=list("abcd")) df[["a", "c", "d"]] ``` vs ```python # R df <- data.frame(a=rnorm(10), b=rnorm(10), c=rnorm(10), d=rnorm(10)) df[, c("a", "c", "d")] ``` ### Data Structures - R - uses data structures like: - Vectors - Lists - Matrices - Dataframes - Pandas - [Intro to data structures](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/dsintro.html?ref=datascientyst.com) - DataFrames - Series DataFrames are the primary data structure for data analysis in R and Pandas. ### Performance R is considered to be faster for most operations in comparison to Pandas. For smaller datasets Pandas might be close to R. To test performance we can use dataset with 2GB/10M rows - [Game Recommendations on Steam](https://www.kaggle.com/datasets/antonkozyriev/game-recommendations-on-steam?select=recommendations.csv&ref=datascientyst.com): ```python # pandas %%time import pandas as pd df = pd.read_csv('recommendations.csv') df['hours'].mean() # R library(microbenchmark) microbenchmark(df <- read.csv('recommendations.csv'), mean(df[, 'hours'])) end ``` The results are: - Pandas ```python CPU times: user 17.1 s, sys: 4.38 s, total: 21.5 s Wall time: 23.7 s 103.97299330788391 ``` - R timing | expr | min | lq | mean | median | uq | max neval | | | --------------------- | ---- | ---- | ---- | ------ | ---- | --------- | -- | | df <- read.csv | 141 | 141 | 142 | 141 | 142 | 143 | 10 | | mean(df\[, "hours"\]) | 0.11 | 0.11 | 0.11 | 0.11 | 0.11 | 0.11 | 10 | As we can see times are close for R and Pandas for this use case. ### Package Ecosystem Both offer mature package systems with a wide variety of packages related to data analysis and visualization. - R has a vast repository of packages on CRAN (Comprehensive R Archive Network) dedicated to statistics, data analysis, and visualization. - Pandas is part of the Python ecosystem, which has a broader range of packages for various purposes beyond data analysis. ### Community - R has a strong community of experienced statisticians and data analysts, and there are numerous resources and documentation available for R users. - Pandas benefits from the larger Python community, which offers extensive resources and documentation for data analysis and programming in general. People from different scientific areas join Python and Pandas communities to solve everyday problems. ### Learning Curve Again it depends on personal choice. Python is considered as one of the best programming languages for beginners. R is far below Python in recent surveys for loved language: [stackoverflow survey - Most loved, dreaded, and wanted](https://survey.stackoverflow.co/2022/?ref=datascientyst.com#technology-most-loved-dreaded-and-wanted) ## 3\. Pandas vs R - useful links | | Pandas | R | | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | | data analysis tool | language for statistical computing | | site | [https://pandas.pydata.org/](https://pandas.pydata.org/?ref=datascientyst.com) | [https://www.r-project.org/](https://www.r-project.org/?ref=datascientyst.com) | | docs | [https://pandas.pydata.org/docs/](https://pandas.pydata.org/docs/?ref=datascientyst.com) | [https://cran.r-project.org/manuals.html](https://cran.r-project.org/manuals.html?ref=datascientyst.com) | | packages | [https://pypi.org/](https://pypi.org/?ref=datascientyst.com) | [https://cran.r-project.org/web/packages/](https://cran.r-project.org/web/packages/?ref=datascientyst.com) | | repo | [https://github.com/pandas-dev/pandas](https://github.com/pandas-dev/pandas?ref=datascientyst.com) | \- | | cheatsheet | [Data Wrangling with pandas](https://pandas.pydata.org/Pandas%5FCheat%5FSheet.pdf?ref=datascientyst.com) | [Data Wrangling with dplyr and tidyr](https://www.rstudio.com/wp-content/uploads/2015/02/data-wrangling-cheatsheet.pdf?ref=datascientyst.com) | | basics | [https://pandas.pydata.org/docs/user\_guide/basics.html](https://pandas.pydata.org/docs/user%5Fguide/basics.html?ref=datascientyst.com) | [https://cran.r-project.org/doc/manuals/r-release/R-intro.pdf](https://cran.r-project.org/doc/manuals/r-release/R-intro.pdf?ref=datascientyst.com) | | getting started | [https://pandas.pydata.org/docs/getting\_started/index.html](https://pandas.pydata.org/docs/getting%5Fstarted/index.html?ref=datascientyst.com) | [https://education.rstudio.com/learn/beginner/](https://education.rstudio.com/learn/beginner/?ref=datascientyst.com) | | indexing | 0 based | 1 based | | missing value | np.nan | NA | | Boolean | False/True | FALSE/TRUE | | Comments | \# comment | \# comment | ## 4\. Summary & Resources In summary, Pandas and R are both powerful tools for data analysis, visualization and manipulation. Ultimately, the choice between R and Pandas often depends on your specific needs, existing familiarity with a programming language, and the ecosystem of packages that best suit your data analysis tasks. Personally I find Pandas easier to learn and start because of the previous experience in Python language. Knowing Pandas or R makes it easier to transition to the other one. - [Comparison with R / R libraries](https://pandas.pydata.org/docs/getting%5Fstarted/comparison/comparison%5Fwith%5Fr.html?ref=datascientyst.com) - [Pandas Cheat Sheet for Data Science](https://datascientyst.com/pandas-cheat-sheet-for-data-science/) - [Pandas vs SQL Cheat Sheet](https://datascientyst.com/pandas-vs-sql-cheat-sheet/) - [Pandas vs Julia - cheat sheet and comparison](https://datascientyst.com/pandas-vs-julia-comparison-cheat-sheet/) - [pandas notebook](https://github.com/softhints/Pandas-Exercises-Projects/blob/main/cheat%5Fsheet/pandas%5Fjulia/pandas.ipynb?ref=datascientyst.com) ## 5\. Pandas vs R Cheat Sheet Image Dark version: ![](https://datascientyst.com/content/images/2023/12/Pandas-vs-R-dark.webp) Light Version: ![Pandas vs R light.webp](https://datascientyst.com/content/images/2023/12/Pandas-vs-R-light.webp) ## 6\. Pandas vs R comparison We are working on a visual comparison between R and Pandas. Below you can find a quick teaser: ![pandas vs R comparison.webp](https://datascientyst.com/content/images/2023/12/pandas-vs-R-comparison.webp) P.S. We were overloaded in the last year so we were not able to post frequently. We hope to have more time for this project and data science. ### Trim Leading & Trailing White Space in Pandas DataFrame URL: https://datascientyst.com/trim-leading-trailing-white-space-pandas-dataframe/ Last updated: 2023-11-24T14:09:13.000Z **To trim leading and trailing whitespaces from strings in Pandas DataFrame**, you can use the `str.strip()` function to trim leading and trailing whitespaces from strings. Here's an example: ```python import pandas as pd data = {'col_1': [' Apple ', ' Banana', 'Orange ', ' Grape '], 'col_2': [' Red ', 'Yellow ', ' Orange', 'Purple ']} df = pd.DataFrame(data) ``` sample data is: | | col\_1 | col\_2 | | - | ------ | ------ | | 0 | Apple | Red | | 1 | Banana | Yellow | | 2 | Orange | Orange | | 3 | Grape | Purple | Browsers usually hide extract spaces at start and end - so you can find row DataFrame data below: ``` col_1 col_2 0 Apple Red 1 Banana Yellow 2 Orange Orange 3 Grape Purple ``` ## 1\. remove spaces - whole DataFrame **To remove leading and trailing whitespaces** in all columns we can: ```python df = df.applymap(lambda x: x.strip() if isinstance(x, str) else x) ``` result: ``` col_1 col_2 0 Apple Red 1 Banana Yellow 2 Orange Orange 3 Grape Purple ``` In this example, the `applymap()` function is used to apply the `strip()` method to each element in the DataFrame. The lambda function checks if the element is a string (isinstance(x, str)) before applying the strip method to avoid errors for non-string elements. After running this code, you'll get a DataFrame where leading and trailing whitespaces in all string columns have been removed. ## 2\. strip spaces in single column To remove trailing and leading spaces from a single column in Pandas DataFrame we can use: - Series method `strip()` - `map` \+ `strip()`: ```python df['col_1'].str.strip() ``` ```python df['col_1'].map(lambda x: x.strip() if isinstance(x, str) else x) ``` or output: ``` 0 Apple 1 Banana 2 Orange 3 Grape Name: col_1, dtype: object ``` ## 3\. trim trailing spaces with regex As alternative solution we an use method `replace()` to remove leading and trailing whitespaces. ```python df.replace(r"^ +| +$", r"", regex=True) ``` result: ``` col_1 col_2 0 Apple Red 1 Banana Yellow 2 Orange Orange 3 Grape Purple ``` ## lstrip + rstrip Pandas offers to handy methods for removing extra spaces: `rstrip` and `lstrip`: ```python s.str.rtrip() s.str.ltrip() ``` ![](https://datascientyst.com/content/images/2023/11/trim-leading-trailing-white-space-pandas-dataframe.png) ## Resources - [pandas.Series.str.strip](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.strip.html?ref=datascientyst.com) - [pandas.Series.str.rstrip](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.rstrip.html?ref=datascientyst.com) - [pandas.Series.str.lstrip](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.lstrip.html?ref=datascientyst.com) - [pandas.Series.replace](https://pandas.pydata.org/docs/reference/api/pandas.Series.replace.html?ref=datascientyst.com) - [pandas.DataFrame.replace](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.replace.html?ref=datascientyst.com) - [Pandas trim leading & trailing white space in a dataframe](https://stackoverflow.com/questions/49551336/pandas-trim-leading-trailing-white-space-in-a-dataframe?ref=datascientyst.com) ### How To Map DataFrame Index to Dictionary in Pandas URL: https://datascientyst.com/map-dataframe-index-to-dictionary-pandas/ Last updated: 2023-11-24T07:53:40.000Z In this post, we'll explore how to **map DataFrame Index values using a dictionary in Pandas**. ## Setup Consider a DataFrame with following data: ```python import pandas as pd data = {'Value': [10, 15, 20, 25]} df = pd.DataFrame(data, index=[1,2,3,4]) ``` result: | | Value | | - | ----- | | 1 | 10 | | 2 | 15 | | 3 | 20 | | 4 | 25 | This will create a DataFrame with an index labeled 1, 2, 3 and 4. ## 1: Map index with `df.index.map` To map DataFrame index with Python dictionary we can use method: `df.index.map`: ```python index_mapping = {1: 'Red', 2: 'Blue', 3: 'Green', 4: 'White' } df.index = df.index.map(index_mapping) ``` The new index is based on the mapping of the provided values in the dictionary: | | Value | | ----- | ----- | | Red | 10 | | Blue | 15 | | Green | 20 | | White | 25 | ## 2: Map with a function To map Pandas index with a function we have two options: - lambda - predefined functions ### lambda Let's remind us that - lambda function is a small anonymous function. ```python df.index.map(lambda x: x + 1) ``` the result is new index with changed values: ``` Index([2, 3, 4, 5], dtype='int64') ``` Another lambda example to map index: ```python df.index.map(lambda x: x.upper()) ``` ### predefined functions The example below will map all values and format the them: ```python df.index.map('Index {}'.format) ``` the result is new index with changed values: ``` Index(['Index 1', 'Index 2', 'Index 3', 'Index 4'], dtype='object') ``` ## 3: Missing values in the dict There is a parameter `na_action` which controls behavior of missing values. If the index contains missing values they could be exclude from mapping with function: ```python import pandas as pd data = {'Value': [10, 15, 20, 25]} df = pd.DataFrame(data, index=[1,2,3, None]) df.index.map('Index {}'.format, na_action='ignore') ``` Will result into: ``` Index(['Index 1.0', 'Index 2.0', 'Index 3.0', nan], dtype='object') ``` ## 4: Map with values without mapping If a value is not found in the index we will end with index full of NaN values: ```python index_mapping = {1: 'Red', 2: 'Blue'} df.index = df.index.map(index_mapping) ``` result: | | Value | | ---- | ----- | | Red | 10 | | Blue | 15 | | NaN | 20 | | NaN | 25 | To avoid that we can use method replace: ```python index_mapping = {1: 'Red', 2: 'Blue'} df.index = pd.Series(df.index).replace(index_mapping) ``` ## Conclusion Mapping a DataFrame index using a dictionary might help to control index values. Some use cases are: - data cleaning - memory efficiency - anonymization ![](https://datascientyst.com/content/images/2023/11/map-dataframe-index-to-dictionary-pandas.png) ## Resources - [pandas.Index.map](https://pandas.pydata.org/docs/reference/api/pandas.Index.map.html?ref=datascientyst.com) - [Index objects](https://pandas.pydata.org/docs/reference/indexing.html?ref=datascientyst.com) - [Python 'map' function inserting NaN, possible to return original values instead?](https://stackoverflow.com/questions/35589820/python-map-function-inserting-nan-possible-to-return-original-values-instead?ref=datascientyst.com) - [In pandas, what does the na\_action parameter to Series.map do](https://stackoverflow.com/questions/39461328/in-pandas-what-does-the-na-action-parameter-to-series-map-do?ref=datascientyst.com) - [Map dataframe index using dictionary](https://stackoverflow.com/questions/43356704/map-dataframe-index-using-dictionary?ref=datascientyst.com) ### Pandas read_csv: Automatic Date Reading from CSV Files URL: https://datascientyst.com/pandas-read_csv-automatic-date-reading-from-csv-files/ Last updated: 2023-11-22T00:04:35.000Z In this article, we will see how **Pandas handles dates during the CSV reading process and automatic date recognition with method `read_csv()`**. ## Automatic Date Reading in Pandas Pandas is designed to automatically recognize and parse dates while reading data from a CSV file, provided that dates are formatted consistently and we provide details about them. The library uses several parameters: - `parse_dates` - `date_format` - `date_parser` in its `read_csv()` function to enable automatic date parsing. You can read more about this parameter here: [pandas.read\_csv](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) The behavior of parameter `parse_dates` is as follows: - `bool`. If `True` \-> try parsing the index. - `list` of `int` or names. e.g. If `[1, 2, 3]` \-> try parsing columns 1, 2, 3 each as a separate date column. - `list` of `list`. e.g. If `[[1, 3]]` \-> combine columns 1 and 3 and parse as a single date column. - `dict`, e.g. `{'foo' : [1, 3]}` \-> parse columns 1, 3 as date and call result ‘foo’ ## Setup Suppose we work with the following CSV file: ``` Date,Value 2023-01-01,5 2023-01-02,15 2023-01-03,25 ``` or: ``` Date,Value,Time 2023-01-01,5,5:45 2023-01-02,15,6:17 2023-01-03,25,8:20 ``` ## Read Date Columns By default Pandas will not parse date columns. We need to set which columns to be parsed as dates: ```python import pandas as pd df = pd.read_csv('data.csv', parse_dates=['Date']) ``` This dataframe will have Date columns which are of type `datetime64[ns]`. The read dataframe: ``` Date Value 0 2023-01-01 5 1 2023-01-02 15 2 2023-01-03 25 ``` ## read\_csv + date\_format + parse\_dates After Pandas 2.0 we can apply custom formatting by using parameter `date_format`: ```python import pandas as pd df = pd.read_csv('data.csv', parse_dates=['Date'], date_format={'Date': '%Y-%m-%d'}) ``` ## Automatic Index Date Parsing To parse datetime index in Pandas while reading CSV file we can use: - `parse_dates=True` - `index_col='Date'` Example: ```python import pandas as pd df = pd.read_csv('data.csv', parse_dates=True, index_col='Date') ``` The index will look like: ``` DatetimeIndex(['2023-01-01', '2023-01-02', '2023-01-03'], dtype='datetime64[ns]', name='Date', freq=None) ``` and data is: ``` Value Date 2023-01-01 5 2023-01-02 15 2023-01-03 25 ``` ## Custom Function with read\_csv() We can force custom date parsing with custom function in Pandas by `date_parser`: ```python import pandas as pd custom_date_parser = lambda x: pd.to_datetime(x, format='%Y-%d-%m') df = pd.read_csv('data.csv', parse_dates=['Date'], date_parser=custom_date_parser) ``` We are defining a custom date parsing function and then use it by parameter `date_parser`. DataFrame will be read as: ``` Date Value 0 2023-01-01 5 1 2023-02-01 15 2 2023-03-01 25 ``` ## Combine Two columns into single datetime With Pandas we can read separate columns - date and time into a single datetime column by using parameter - `parse_dates`: ```python df = pd.read_csv('data/data.csv', parse_dates={'datetime': ['Date', 'Time']}) ``` Parameter `parse_dates` takes a mapping of the expected type and the columns: `{'datetime': ['Date', 'Time']}` ## Conclusion Pandas offers several ways to parse dates from CSV files. **Pandas provides the flexibility to handle various date formats by automatic date recognition or custom parsing**. Converting dates during `read_csv` operation might prevent errors and offer additional functionality on those columns. ### How to Convert Column to Categorical in Pandas DataFrame with Examples URL: https://datascientyst.com/convert-column-to-categorical-pandas-dataframe-examples/ Last updated: 2023-11-20T23:09:59.000Z In this article, we'll explore how to **convert columns to categorical in a Pandas DataFrame** with practical examples. In data analysis, efficient memory usage and improved performance are crucial considerations. Conversion column to categorical is simple as: ```python df['col'].astype('category') ``` Let's dive into more details. ## Why Use Categorical Data Type? **Categorical data** type is beneficial when dealing with columns containing a limited and fixed set of unique values. This not only **optimizes memory usage** but also **enhances the performance of certain operations**, such as groupby and value\_counts. **Note:** Categorical columns can save up to 50 - 80% memory. ## Examples 1: Basic Conversion to Categorical Consider a DataFrame with a column representing different car types. We can convert this column to a categorical type using the astype method: ```python import pandas as pd data = {'CarType': ['Sedan', 'SUV', 'Truck', 'Sedan', 'Truck']} df = pd.DataFrame(data) df['CarType'] = df['CarType'].astype('category') ``` ## Example 2: Specifying Categories and Order You can specify custom categories and their order using the pd.Categorical constructor. Let's consider a DataFrame with a 'Size' column: ```python data = {'Size': ['Small', 'Medium', 'Large', 'Small']} df = pd.DataFrame(data) df['Size'] = pd.Categorical(df['Size'], categories=['Small', 'Medium', 'Large'], ordered=True) print(df['Size'].cat.categories) ``` ## Check if column is Categorical To check if a column is categorical we can use: ```python df.dtypes ``` You can find the difference before and after conversion: - before - CarType object - converted to categorical - CarType category ## Check which columns are good for Categorical There are different ways to find out if a column is suitable for a categorical. Such column contain: - limited - fixed set of unique values. We can use methods like: ```python df['col'].value_counts() df.describe(how='all') pd.get_dummies(s) ``` To find potential columns for conversion. ## Test memory usage gain To test what is the benefit of using categorical columns in Pandas we will run: ```python import pandas as pd data = {'CarType': ['Sedan', 'SUV', 'Truck', 'Sedan', 'Truck'] * 10000} df = pd.DataFrame(data) ``` We have two option to find memory usage in Pandas: ```python df.memory_usage(deep=True) df.info(memory_usage='deep') ``` sample output: ``` RangeIndex: 50000 entries, 0 to 49999 Data columns (total 1 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 CarType 50000 non-null category dtypes: category(1) memory usage: 49.2 KB ``` Below you can compare the results before and after conversion: - before - memory usage: 2.9 MB - after - memory usage: 49.2 KB This simple example shows the great benefit of using categorical columns in Pandas. ## Summary **Converting column types to categorical in Pandas** is a powerful technique for **optimizing memory usage and enhancing data analysis performance**. Whether dealing with nominal or ordinal categorical data, Pandas provides versatile tools for customization and conversion. By incorporating these examples into your data analysis workflow, you can leverage the benefits of categorical data types and efficiently handle large datasets. **Note:** Nominal data involves categorization without any ranking, while ordinal data involves both categorization and ranking. ## Resources - [Categorical data in Pandas](https://pandas.pydata.org/docs/user%5Fguide/categorical.html?ref=datascientyst.com) - [How can I dynamically distinguish between categorical data and numerical data?](https://datascience.stackexchange.com/questions/9892/how-can-i-dynamically-distinguish-between-categorical-data-and-numerical-data?ref=datascientyst.com) - [pandas.get\_dummies](https://pandas.pydata.org/docs/reference/api/pandas.get%5Fdummies.html?ref=datascientyst.com) ### How to Deal With Whitespace and Irregular Separators in Pandas Read CSV? URL: https://datascientyst.com/make-separator-in-pandas-read_csv-more-flexible-wrt-whitespace-for-irregular-separators/ Last updated: 2023-09-05T12:51:17.000Z In this post, we will try to address how to deal with: - consecutive whitespaces as delimiters - irregular Separators while reading a CSV file with Pandas `read_csv()` method. ## Steps work with irregular separators - Inspect the CSV file - Select Pandas method: - `read_csv` - `read_txt` - `read_fwf` - Test reading with different parameters - `skipinitialspace` - `engine` - `delim_whitespace` - `escapechar` - Select separator - normal one or regex - `sep='\s+'` - `sep='\r\t'` ## Data Suppose we would like to read the following csv file into Pandas DataFrame: "data.csv": ``` Date Company A Company A Company B Company B 2021-09-06 1 7.9 2 6 2021-09-07 1 8.5 2 7 2021-09-08 2 8 1 8.1 ``` ## Example - regex separator - sep='\\s{3}' To correctly read data into DataFrame we need to use a combination of arguments: `sep='\s{3}', engine='python'`. This is needed because we have exactly 3 spaces as delimiter: ```python import pandas as pd df = pd.read_csv('data.csv', sep='\s{3}', engine='python') ``` The resulted DataFrame is: | | Date | Company A | Company A.1 | Company B | Company B.1 | | - | ---------- | --------- | ----------- | --------- | ----------- | | 0 | 2021-09-06 | 1 | 7.9 | 2 | 6.0 | | 1 | 2021-09-07 | 1 | 8.5 | 2 | 7.0 | | 2 | 2021-09-08 | 2 | 8.0 | 1 | 8.1 | **Note**: If we try to use: - `pd.read_csv('data.csv')` \- return single column - `pd.read_csv('data.csv', sep='\s+')` \- column names contain spaces - `pd.read_csv('data.csv', sep='\s*')` \- column names contain spaces all of them will return wrong columns. ## Example - regex separator - two and more spaces We can also define use a regex to define separators as follow: two and more spaces: `sep=r"[ ]{2,}"` ```python import pandas as pd df = pd.read_csv('data.csv', sep='[ ]{2,}', engine='python') ``` Alternative solution might be: `sep='\t\s+'` combination of tabs and spaces as separator in Pandas. **Note**: Why do we need engine='python'? This is explained from the warning: > Falling back to the 'python' engine because the 'c' engine does not support regex separators ## Additional notes: - `delim_whitespace` is equivalent to `sep='\s+'`. If True `sep` can be skipped - `skipinitialspace` \- skip spaces after delimiter - Separators longer than 1 character and different from '\\s+' - will be interpreted as regular expressions - also force the use of the Python parsing engine. - Always analyze and clean spaces because can cause issues: - "hello " != " hello" - how values of empty spaces should be treated - errors, missing values etc? ## Output ![](https://datascientyst.com/content/images/2023/09/make-separator-in-pandas-read_csv-more-flexible-wrt-whitespace-for-irregular-separators.webp) ## Resources - [How to Use Multiple Char Separator in read\_csv in Pandas](https://datascientyst.com/use-multiple-char-separator-read%5Fcsv-pandas/) - [How to Read Data from Text File Into Pandas](https://datascientyst.com/load-data-from-text-file-into-pandas/) - [pandas.read\_csv](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) ### How to split dataframe in Pandas URL: https://datascientyst.com/how-to-split-dataframe-in-pandas/ Last updated: 2023-09-04T20:46:22.000Z In this short guide, I'll **show you how to split Pandas DataFrame**. You can also find how to: - split a large Pandas DataFrame - pandas split dataframe into equal chunks - split DataFrame by percentage - split dataset into training and testing parts To start, here is the syntax to split Pandas Dataframe in 5 equal chunks: ```python import numpy as np np.array_split(df, 5) ``` which returns a list of DataFrames. Let's see all the steps in details ## Setup Lets create a sample DataFrame which contains 12 rows: ```python import pandas as pd import numpy as np data = np.random.randint(0,12,size=(12, 4)) df = pd.DataFrame(data, columns=list('ABCD')) ``` data looks like: | | A | B | C | D | | -- | -- | -- | -- | - | | 0 | 8 | 4 | 7 | 4 | | 1 | 1 | 9 | 4 | 5 | | 2 | 1 | 4 | 1 | 8 | | 3 | 4 | 9 | 3 | 7 | | 4 | 8 | 11 | 5 | 9 | | 5 | 11 | 9 | 11 | 3 | | 6 | 5 | 1 | 6 | 3 | | 7 | 2 | 11 | 11 | 7 | | 8 | 5 | 9 | 4 | 4 | | 9 | 9 | 0 | 11 | 0 | | 10 | 4 | 10 | 3 | 8 | | 11 | 0 | 0 | 2 | 1 | ## Step 1: Split dataframe into n chunks Numpy method `np.array_split()` can be used on Pandas DataFrame to split it in n chunks: ```python import pandas as pd import numpy as np df = pd.DataFrame(np.random.randint(0,12,size=(12, 4)), columns=list('ABCD')) chunks = np.array_split(df, 5) for chunk in chunks: print(chunk.shape) display(chunk) ``` The result is 5 chunks. As you can see the first two chunks has 3 rows while the rest 2 rows: ![](https://datascientyst.com/content/images/2023/05/split-pandas-dataframe.webp) ## Step 2: Split DataFrame with list comprehension To split DataFrame by using list comprehensions we can: - calculate the chunk size - get all rows for a given rage ```python chunk_size = df.shape[0] // 3 chunks = [df[i:i+chunk_size].copy() for i in range(0, df.shape[0], chunk_size)] ``` This will divide the input DataFrame into 3 DataFrames: ``` [ A B C D 0 8 4 7 4 1 1 9 4 5 2 1 4 1 8 3 4 9 3 7, A B C D 4 8 11 5 9 5 11 9 11 3 6 5 1 6 3 7 2 11 11 7, A B C D 8 5 9 4 4 9 9 0 11 0 10 4 10 3 8 11 0 0 2 1] ``` ## Step 3: Split DataFrame into groups We can use method `groupby` to split DataFrame into: - equal chunks - group by criteria The next code will split DataFrame into equal sized groups: ```python groups = df.groupby(df.index % 4) ``` Now we can work with each group to get information like count or sum: ```python groups.count() groups.sum() ``` result: | | A | B | C | D | | - | - | - | - | - | | 0 | 3 | 3 | 3 | 3 | | 1 | 3 | 3 | 3 | 3 | | 2 | 2 | 2 | 2 | 2 | | 3 | 2 | 2 | 2 | 2 | | 4 | 2 | 2 | 2 | 2 | ## Step 4: Split DataFrame by column We can group by a column and then split the DataFrame: ```python groups = df.groupby(df['A'] % 2) ``` To get information for a group we can use method `first`: ```python groups.first() ``` or display all groups ```python for group in groups: display(group) ``` result: ``` (0, A B C D 0 8 4 7 4 5 11 9 11 3 10 4 10 3 8) (1, A B C D 1 1 9 4 5 6 5 1 6 3 11 0 0 2 1) (2, A B C D 2 1 4 1 8 7 2 11 11 7) (3, A B C D 3 4 9 3 7 8 5 9 4 4) (4, A B C D 4 8 11 5 9 9 9 0 11 0) ``` ## Step 5: Split large DataFrame To split large DataFrames we can use the Dask library. We can: - convert the Pandas DataFrame to Dask - provide number of partitions To split large DataFrame with Dask we can do: ### Read CSV file with Dask and split to DataFrames In case that you need to read huge CSV file and split it to several DataFrames you can use Dask as follow: ```python import dask.dataframe as dd df = dd.read_csv("data.csv").repartition(npartitions=3) ``` ### Split large DataFrame with Dask If the DataFrame exists it can be split to chunks by this method `repartition`: ```python df = df.repartition(npartitions=2) ``` ## Step 6: Split dataframe by percentage Finally let say that you need to split your DataFrame to several parts taking into account percentage. For this purpose we can use `train_test_split` from `sklearn`. We can specify the training and test size as a percentage. ### Split to train and testing data This is escpecially good for machine learning when you need to split DataFrame into 2 parts for testing and training: ```python from sklearn.model_selection import train_test_split x, x_test, y, y_test = train_test_split(df,df.index,test_size=0.2,train_size=0.8) ``` where: - x - contains the first DataFrame - y - indexes of first DataFrame - Index(\[0, 3, 6, 9, 2, 8, 10, 11, 7\], dtype='int64') - x\_test - second DataFrame - y\_test - indexes of the second one - Index(\[4, 5, 1\], dtype='int64') ### Split by percentage with numpy We can also use numpy to split the DataFrame into multiple datasets based on percentage from the original data. This can be done by: ```python partitions = [int(.2*len(df)), int(.5*len(df)), int(.6*len(df))] a, b, c, d = np.split(df, partitions) ``` So partitions are calculated and defined as: ``` [2, 6, 7] ``` which selects: - 2 rows for the first one - 4 rows for the second - 1 rows for the 3rd one - the rest for the 4th ## Conclusion In this post we saw multiple different ways to split DataFrame into several chunks. We have used different additional packages like `numpy`, `sklearn` and `dask`. You will know how to easily split DataFrame into training and testing datasets. We also covered how to read a huge CSV file and separate it into multiple DataFrames with Dask. ## Resources - [numpy.split](https://numpy.org/doc/stable/reference/generated/numpy.split.html?ref=datascientyst.com) - [sklearn.model\_selection .train\_test\_split](https://scikit-learn.org/stable/modules/generated/sklearn.model%5Fselection.train%5Ftest%5Fsplit.html?ref=datascientyst.com) - [dask.dataframe.DataFrame.repartition](https://docs.dask.org/en/stable/generated/dask.dataframe.DataFrame.repartition.html?ref=datascientyst.com) ### error: nothing to repeat at position 0 - Pandas URL: https://datascientyst.com/error-nothing-to-repeat-at-position-0-pandas/ Last updated: 2023-09-04T14:18:09.000Z Common Pandas error - error: nothing to repeat at position 0 can be result of several operations: - bad regular expression - reading CSV file with incorrect separator ## Fix error: nothing to repeat at position 0 To fix error: > error: nothing to repeat at position 0 First we need to identify what is the reason for this error. Once we know the reason the solutions might be: - change - `.str.contains('?')` to `.str.contains('?', regex=False)` - change the `read_csv()` sepators, encoding etc: - `, sep='\t', engine='python', encoding='ISO-8859-1'` ## Data Suppose we have data like the one below: ```python from faker import Faker import pandas as pd Faker.seed(0) fake = Faker() addr = [] for _ in range(5): addr.append(fake.address()) df = pd.DataFrame({'address':addr}) ``` | | address | | - | ---------------------------------------------------- | | 0 | 48764 Howard Forge Apt. 421\\nVanessaside, VT 79393 | | 1 | PSC 4115, Box 7815\\nAPO AA 41945 | | 2 | 778 Brown Plaza\\nNorth Jenniferfurt, VT 88077 | | 3 | 3513 John Divide Suite 115\\nRodriguezside, LA 93111 | | 4 | 398 Wallace Ranch Suite 593\\nIvanburgh, AZ 80818 | ## Example - bad regex If we try to search for a question mark we will end with Pandas error: **error: nothing to repeat at position 0.** ```python df['address'].str.contains('?') ``` The solution is either to escape the question mark or use `regex=False`: ```python df['address'].str.contains('\?') df['address'].str.contains('?', regex=False) ``` ## Fix - read\_csv To fix the same error if we get it during reading CSV file by method `read_csv` file we can use parameters like: ```python pd.read_csv('data.csv', sep=''\*,\*'' , engine='python', encoding='ISO-8859-1') ``` Why do we get this error for read\_csv? The reason is that some characters are special and treated in a different way. So we may need to escape them: > addition, separators longer than 1 character and different from '\\s+' will be interpreted as regular expressions and will also force the use of the Python parsing engine. Note that regex delimiters are prone to ignoring quoted data. Regex example: '\\r\\t'. ## Output ![](https://datascientyst.com/content/images/2023/09/error-nothing-to-repeat-at-position-0-pandas.webp) ## Resources - [pandas.read\_csv](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) - [pandas.Series.str.contains](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.contains.html?ref=datascientyst.com) ### ValueError: pattern contains no capture groups - pandas URL: https://datascientyst.com/valueerror-pattern-contains-no-capture-groups-pandas/ Last updated: 2023-09-04T10:06:19.000Z To **solve Pandas Error: Valueerror: Pattern Contains No Capture Groups we need to specify a capture group**. ## Steps to plot 2 variables - Import matplotlib library - Create DataFrame with correlated data - Create the figure and axes object - `fig, ax = plt.subplots()` - Plot the first variable on x and left y axes - Plot the second variable on x and secondary y axes More information can be found: [DataFrame.plot - secondary\_y](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.plot.html?ref=datascientyst.com) ## Data Let's have this DataFrame which will be used to demonstrate the error: > Valueerror: Pattern Contains No Capture Groups ```python from faker import Faker import pandas as pd Faker.seed(0) fake = Faker() addr = [] for _ in range(5): addr.append(fake.address()) df = pd.DataFrame({'address':addr}) df ``` | | address | | - | ---------------------------------------------------- | | 0 | 48764 Howard Forge Apt. 421\\nVanessaside, VT 79393 | | 1 | PSC 4115, Box 7815\\nAPO AA 41945 | | 2 | 778 Brown Plaza\\nNorth Jenniferfurt, VT 88077 | | 3 | 3513 John Divide Suite 115\\nRodriguezside, LA 93111 | | 4 | 398 Wallace Ranch Suite 593\\nIvanburgh, AZ 80818 | ## Example The code below produce the error: ```python df['address'].str.extract('.+?(?=\n)') ``` to solve the error we will add parentheses to denote the capture group: ```python df['address'].str.extract('(.+)?(?=\n)') ``` So the fix is: - `'.+?(?=\n)'` - `'(.+)?(?=\n)'` which produce: | | 0 | | - | --------------------------- | | 0 | 48764 Howard Forge Apt. 421 | | 1 | PSC 4115, Box 7815 | | 2 | 778 Brown Plaza | | 3 | 3513 John Divide Suite 115 | | 4 | 398 Wallace Ranch Suite 593 | ## Output ![](https://datascientyst.com/content/images/2023/09/valueerror-pattern-contains-no-capture-groups-pandas.webp) ## Resources - [pandas.Series.str.extract](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.extract.html?ref=datascientyst.com) - [Grouping in Python Regex](https://docs.python.org/3/howto/regex.html?ref=datascientyst.com#grouping) - [Non-capturing and Named Groups](https://docs.python.org/3/howto/regex.html?ref=datascientyst.com#non-capturing-and-named-groups) ### How to Extract Everything Before or After with Regex in Pandas URL: https://datascientyst.com/extract-everything-before-after-regex-pandas/ Last updated: 2023-09-04T07:16:10.000Z To plot two variables on two sides of Y-axes, we can plot in two steps: - `'(.*?)\n'` - `'.+?(?=\n)'` ## Steps to extract everything until/after Below are the steps which I usually follow for regex extraction in Pandas - analyse the data from which I will extract - clean the data - choose pandas method - `split`, `extract` etc - define regex pattern - create new column(s) ## Data Let's create simple sample DataFrame to be used for regex extraction: ```python from faker import Faker import pandas as pd Faker.seed(0) fake = Faker() addr = [] for _ in range(5): addr.append(fake.address()) df = pd.DataFrame({'address':addr}) ``` | | address | | - | ---------------------------------------------------- | | 0 | 48764 Howard Forge Apt. 421\\nVanessaside, VT 79393 | | 1 | PSC 4115, Box 7815\\nAPO AA 41945 | | 2 | 778 Brown Plaza\\nNorth Jenniferfurt, VT 88077 | | 3 | 3513 John Divide Suite 115\\nRodriguezside, LA 93111 | | 4 | 398 Wallace Ranch Suite 593\\nIvanburgh, AZ 80818 | ## Example 1 - Captcharing group and characters Extract everything in Pandas column up to new line ```python df['address'].str.extract('(.*?)\n') ``` result: | | 0 | | - | --------------------------- | | 0 | 48764 Howard Forge Apt. 421 | | 1 | PSC 4115, Box 7815 | | 2 | 778 Brown Plaza | | 3 | 3513 John Divide Suite 115 | | 4 | 398 Wallace Ranch Suite 593 | ## Example 2 - Non captcharing groups Extract everything in Pandas column up to new line ```python df['address'].str.extract('(.+)?(?=\n)') ``` result: | | 0 | | - | --------------------------- | | 0 | 48764 Howard Forge Apt. 421 | | 1 | PSC 4115, Box 7815 | | 2 | 778 Brown Plaza | | 3 | 3513 John Divide Suite 115 | | 4 | 398 Wallace Ranch Suite 593 | ## Output ![](https://datascientyst.com/content/images/2023/09/extract-everything-before-after-regex-pandas.png) ## Resources - [Working with text data](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/text.html?ref=datascientyst.com) - [pandas.Series.str.extract](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.extract.html?ref=datascientyst.com) - [pandas.Series.str.contains](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.contains.html?ref=datascientyst.com) ### How To Split Column by Multiple Characters with Regex in Pandas URL: https://datascientyst.com/split-column-by-multiple-characters-regex-in-pandas/ Last updated: 2023-09-03T20:52:40.000Z To split Pandas column by multiple characters we can use complex regex pattern as: - `df['address'].str.split('; |, |\n', expand=True)` - `df['address'].str.extract(r'(.*)\n(.*)')` ## Steps to split column in Pandas - Import matplotlib library - Create DataFrame with correlated data - Create the figure and axes object - `fig, ax = plt.subplots()` - Plot the first variable on x and left y axes - Plot the second variable on x and secondary y axes More information can be found: [DataFrame.plot - secondary\_y](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.plot.html?ref=datascientyst.com) ## Data Suppose we have DataFrame with Fake address data: ```python from faker import Faker import pandas as pd Faker.seed(0) fake = Faker() addr = [] for _ in range(5): addr.append(fake.address()) df = pd.DataFrame({'address':addr}) ``` Data should be something like: | | address | | - | ---------------------------------------------------- | | 0 | 48764 Howard Forge Apt. 421\\nVanessaside, VT 79393 | | 1 | PSC 4115, Box 7815\\nAPO AA 41945 | | 2 | 778 Brown Plaza\\nNorth Jenniferfurt, VT 88077 | | 3 | 3513 John Divide Suite 115\\nRodriguezside, LA 93111 | | 4 | 398 Wallace Ranch Suite 593\\nIvanburgh, AZ 80818 | Check resources to find out how to create more fake data with Pandas. ## Example 1 - str.split We can use method `str.split` with parameter `expand=True` to **split Pandas column by multiple separators** and expand content into new columns: ```python df['address'].str.split('; |, |\n', expand=True) ``` output: | | 0 | 1 | 2 | | - | --------------------------- | ------------------ | ------------ | | 0 | 48764 Howard Forge Apt. 421 | Vanessaside | VT 79393 | | 1 | PSC 4115 | Box 7815 | APO AA 41945 | | 2 | 778 Brown Plaza | North Jenniferfurt | VT 88077 | | 3 | 3513 John Divide Suite 115 | Rodriguezside | LA 93111 | | 4 | 398 Wallace Ranch Suite 593 | Ivanburgh | AZ 80818 | ## Example 2 - str.extract As alternative we can use `str.extract` and capturing groups to **split string into columns** as follow: ```python df['address'].str.extract(r'(.*)\n(.*)\,(.*)') ``` output: | | 0 | 1 | 2 | | - | --------------------------- | ------------------ | -------- | | 0 | 48764 Howard Forge Apt. 421 | Vanessaside | VT 79393 | | 1 | NaN | NaN | NaN | | 2 | 778 Brown Plaza | North Jenniferfurt | VT 88077 | | 3 | 3513 John Divide Suite 115 | Rodriguezside | LA 93111 | | 4 | 398 Wallace Ranch Suite 593 | Ivanburgh | AZ 80818 | ## Output ![](https://datascientyst.com/content/images/2023/09/split-column-by-multiple-characters-regex-in-pandas.webp) ## Resources - [How To Make a Fake Data Set in Python and Pandas](https://datascientyst.com/make-fake-data-set-python-pandas/) - [pandas.Series.str.split](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.split.html?ref=datascientyst.com) - [pandas.Series.str.extract](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.extract.html?ref=datascientyst.com) ### How To Read Only Specific Columns in Pandas read CSV URL: https://datascientyst.com/how-to-read-only-specific-columns-in-pandas-read-csv/ Last updated: 2023-09-03T09:43:10.000Z To read only specific columns from CSV file using Pandas read\_csv method we need to use parameter `usecols=fields` ## Steps to read specific columns from CSV file - Import pandas - Define columns to be read - `usecols=fields` \- to list columns to be read - Subset of columns to select, denoted either by column labels or column indices. - `low_memory = True` \- reading in chunks - Internally process the file in chunks, resulting in lower memory use while parsing, but possibly mixed type inference. - `index_col` \- to identify column X as index - Column(s) to use as row label(s), denoted either by column labels or column indices. - `names` and `header` to override the column names. More information can be found: [DataFrame.plot - secondary\_y](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.plot.html?ref=datascientyst.com) ## Data Suppose we have the following CSV file which we like to read with Pandas. We want to read only single column from this CSV file into DataFrame: ```bash ,x,y,z 0,a,e,1 1,b,f,2 2,c,g,3 3,d,i,4 ``` | | x | y | z | | - | - | - | - | | 0 | a | e | 1 | | 1 | b | f | 2 | | 2 | c | g | 3 | | 3 | d | i | 4 | ## Example ```python import pandas as pd cols = ['x', 'y'] df = pd.read_csv('data/data_0.csv', usecols = cols, low_memory = True) ``` ## Output ![](https://datascientyst.com/content/images/2023/09/read-only-specific-columns-in-pandas-read-csv.webp) ## Resources - [pandas.read\_csv](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) ### How to Round To Nearest Hour in Pandas URL: https://datascientyst.com/how-to-round-to-nearest-hour-in-pandas/ Last updated: 2023-08-27T15:01:01.000Z To round to closest hour in Pandas datetime column we can several options: **(1) Round to nearest hour** ```python df['date'].dt.round('H').dt.hour ``` **(2) floor to closest hour** ```python df['date'].dt.floor('h') ``` **(3) ceil to closest hour** ```python df['date'].dt.ceil('h') ``` The image below show the results of 3 options: ![](https://datascientyst.com/content/images/2023/08/how-to-round-to-nearest-hour-in-pandas.png) Let's cover the cases in examples. ## Example Suppose we have DataFrame with data: ```python import pandas as pd dates = ['2023-08-27 17:45', '2023-08-27 19:15', '2023-08-27 20:07', '2023-08-27 23:55',] df = pd.DataFrame({'date': dates}) ``` result: | | date | | - | ------------------- | | 0 | 2023-08-27 17:45:00 | | 1 | 2023-08-27 19:15:00 | | 2 | 2023-08-27 20:07:00 | | 3 | 2023-08-27 23:55:00 | Note: if you need convert string to datetime column by: ```python df['date'] = pd.to_datetime(df['date']) ``` ## Round to closest hour First we will try to round to closest hour: ```python df['date'].dt.round('H').dt.hour ``` This will give us the hour as integer: ``` 0 18 1 19 2 20 3 0 Name: date, dtype: int32 ``` ## Floor to nearest hour Instead of rounding we can floor time to hour. Operation floor means: > rounds a number DOWN to the nearest integer, if necessary, and returns the result. ```python df['date'].dt.floor('h').dt.hour ``` This will give us the hour as integer: ``` 0 17 1 19 2 20 3 23 Name: date, dtype: int32 ``` As you can notice two results differ from the first case. ## Ceil to nearest hour Instead of rounding we can also ceil time to hour. Operation ceil means: > rounds a number UP to the nearest integer, if necessary, and returns the result. ```python df['date'].dt.ceil('h').dt.hour ``` This will give us the hour as integer: ``` 0 18 1 20 2 21 3 0 Name: date, dtype: int32 ``` Again we have different results. ## Resources - [pandas.Series.dt.floor](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.floor.html?ref=datascientyst.com) - [pandas.Series.dt.round](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.round.html?ref=datascientyst.com) - [pandas.Series.dt.ceil](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.ceil.html?ref=datascientyst.com) - [How to Round Time to the Nearest Quarter or Hour in Pandas?](https://datascientyst.com/how-to-round-time-to-the-nearest-quarter-or-hour-in-pandas/) ### Pandas pivot_table Silently Drops Indices with NaNs URL: https://datascientyst.com/pandas-pivot_table-silently-drops-indices-with-nans/ Last updated: 2023-08-27T14:38:26.000Z In this post, we will discuss when pivot\_table silently drops indices with NaN-s. We will give an example, expected behavior and many resources. ## Example Let's have a DataFrame like: ```python import pandas as pd import numpy as np df = pd.DataFrame({'foo': ['one', 'one', 'one', 'two', 'two', 'two'], 'bar': ['A', 'B', np.nan, 'A', 'B', 'C'], 'baz': [1, 2, 3, 4, 5, 6], 'zoo': ['x', 'y', 'z', 'q', 'w', 't']}) ``` with data: | | foo | bar | baz | zoo | | - | --- | --- | --- | --- | | 0 | one | A | 1 | x | | 1 | one | B | 2 | y | | 2 | one | NaN | 3 | z | | 3 | two | A | 4 | q | | 4 | two | B | 5 | w | ## silent drop of NaN-s indexes Now let's run two different examples: ### pivot\_table ```python df.pivot_table(index='foo', columns='bar', values='zoo', aggfunc=sum) ``` result is: | bar | A | B | C | | --- | - | - | --- | | foo | | | | | one | x | y | NaN | | two | q | w | t | Even trying with `dropna=False` still results in the same behavior in pandas 2.0.1: ```python df.pivot_table(index='foo', columns='bar', values='zoo', aggfunc=sum, dropna=False) ``` ### pivot\_table and dropna Below you can read what is doing parameter `dropna`: > dropna bool, default True > Do not include columns whose entries are all NaN. If True, rows with a NaN value in any column will be omitted before computing margins. ### pivot while pivot will give us different result: ```python df.pivot(index='foo', columns='bar', values='zoo') ``` which returns NaN-s from the bar column: | bar | nan | A | B | C | | --- | --- | - | - | --- | | foo | | | | | | one | z | x | y | NaN | | two | NaN | q | w | t | ![](https://datascientyst.com/content/images/2023/08/pandas-pivot_table-silently-drops-indices-with-nans.png) ## Stop silent drop Once you analyze the error and data a potential solution might be to fill NaN values with default value (which differs from the rest): ```python df['bar'] = df['bar'].fillna(0) df.pivot_table(index='foo', columns='bar', values='zoo', aggfunc=sum) ``` After this change `pivot_table` will not drop the NaN indexes: | | foo | bar | baz | zoo | | - | --- | --- | --- | --- | | 0 | one | A | 1 | x | | 1 | one | B | 2 | y | | 2 | one | 0 | 3 | z | | 3 | two | A | 4 | q | | 4 | two | B | 5 | w | The image below show the behaviour before and after the silent drop of NaN-s: ![fix-pivot_table-silently-drops-indices-with-nans](https://datascientyst.com/content/images/2023/08/fix-pivot_table-silently-drops-indices-with-nans.png) ## Conclusion You can always refer to the official Pandas documentation for examples and what is expected: [Reshaping and pivot tables](https://pandas.pydata.org/docs/user%5Fguide/reshaping.html?ref=datascientyst.com) Pandas offers a variety of methods and functions to wrangle data. Sometimes the results might be unexpected. In this case test the results against another method or sequence of steps. If you notice a Pandas bug or unexpected behavior you can open ticket or check Pandas issues like: [ENH: pivot/groupby index with nan #3729](https://github.com/pandas-dev/pandas/issues/3729?ref=datascientyst.com) ## Resources - [ENH: pivot/groupby index with nan #3729](https://github.com/pandas-dev/pandas/issues/3729?ref=datascientyst.com) - [DataFrame.pivot](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.pivot.html?ref=datascientyst.com) - [pandas.pivot\_table](https://pandas.pydata.org/docs/reference/api/pandas.pivot%5Ftable.html?ref=datascientyst.com) ### How To Read Multiple CSV Files into Pandas DataFrame URL: https://datascientyst.com/how-to-read-multiple-csv-files-into-pandas-dataframe/ Last updated: 2023-08-27T09:16:10.000Z To read multiple CSV file into single Pandas DataFrame we can use the following syntax: **(1) Pandas read multiple CSV files** ```python path = r'/home/user/Downloads' all_files = glob.glob(path + "/*.csv") lst = [] for filename in all_files: df = pd.read_csv(filename, index_col=None, header=0) lst.append(df) merged_df = pd.concat(lst, axis=0, ignore_index=True) ``` **(2) Read multiple CSV files - Dask** ```python import dask.dataframe as dd df = dd.read_csv("~/Downloads/test*.csv") ``` ## Pandas Example Suppose that we would like to read all CSV files: - located in folder - `/home/user/Downloads` - by pattern - `/test_*.csv` \- starting with `test_` and ending on `.csv` We can use the following code: ```python import glob import pandas as pd path = r'/home/user/Downloads' pattern = "/test_*.csv" all_files = glob.glob(path + pattern) lst = [] for filename in all_files: df = pd.read_csv(filename, index_col=None, header=0) lst.append(df) merged_df = pd.concat(lst, axis=0, ignore_index=True) ``` Let's say that we have the following files in this folder: - other.csv - test\_1.csv - test\_2.csv In the final DataFrame - merged\_df we will have content only from files - test\_1.csv and test\_2.csv: ![](https://datascientyst.com/content/images/2023/08/read-multiple-csv-files-into-pandas-dataframe.webp) ## Read multiple CSV files with Dask As an alternative solution we can use the dask module to read multiple CSV files. To install Dask you can visit: [dask](https://pypi.org/project/dask/?ref=datascientyst.com) or use: `pip install dask`. To read multiple files from a folder with pattern we can use: ```python import dask.dataframe as dd df = dd.read_csv("~/Downloads/test*.csv") ``` ## Resources For more advanced examples on reading multiple CSV or JSON files with Pandas you can check: - [How to Merge multiple CSV Files in Linux Mint](https://softhints.com/merge-multiple-csv-files-linux-mint/?ref=datascientyst.com) - [How to Merge Multiple JSON Files with Python](https://softhints.com/merge-multiple-json-files-pandas-dataframe/?ref=datascientyst.com) - [Convert Pandas to Dask DataFrame ( Dask to Pandas )](https://datascientyst.com/convert-pandas-to-dask-dataframe-dask-to-pandas/) ### How To Margin Only on Single Axis - Column or Row in Pandas URL: https://datascientyst.com/how-to-margin-only-on-single-axis-column-or-row-in-pandas/ Last updated: 2025-01-21T21:23:42.000Z We can use the following syntax to margin on a single axis column or row in Pandas: **(1) Margin only on rows** ```python df.pivot_table(index='foo', columns='bar', values='baz', margins=True).iloc[:, :-1] ``` **(2) Margin only on columns** ```python df.pivot_table(index='foo', columns='bar', values='baz', margins=True).iloc[:-1, :] ``` In the examples above we do margin on both axes but return only a given total by removing results with function `.iloc[:-1, :]`. ![](https://datascientyst.com/content/images/2023/08/how-to-margin-only-on-single-axis-column-or-row-in-pandas.webp) ## Example Suppose we have a DataFrame like: ```python import pandas as pd df = pd.DataFrame({'foo': ['one', 'one', 'one', 'two', 'two', 'two'], 'bar': ['A', 'B', 'B', 'A', 'B', 'C'], 'baz': [1, 2, 3, 4, 5, 6], 'zoo': ['x', 'y', 'z', 'q', 'w', 't']}) ``` with data: | | foo | bar | baz | zoo | | - | --- | --- | --- | --- | | 0 | one | A | 1 | x | | 1 | one | B | 2 | y | | 2 | one | B | 3 | z | | 3 | two | A | 4 | q | | 4 | two | B | 5 | w | To get totals or margins per column or rows we need to pass argument - `margins=True`: ```python df.pivot_table(index='foo', columns='bar', values='baz', margins=True) ``` result: | bar | A | B | C | All | | --- | --- | --- | --- | --- | | foo | | | | | | one | 1.0 | 2.0 | 3.0 | 2.0 | | two | 4.0 | 5.0 | 6.0 | 5.0 | | All | 2.5 | 3.5 | 4.5 | 3.5 | ## 1\. Margin only on rows To margin only on rows we can do: ```python df.pivot_table(index='foo', columns='bar', values='baz', margins=True).iloc[:, :-1] ``` Which results into new row with totals: | bar | A | B | C | | --- | --- | --- | --- | | foo | | | | | one | 1.0 | 2.0 | 3.0 | | two | 4.0 | 5.0 | 6.0 | | All | 2.5 | 3.5 | 4.5 | ## 2\. Margin only on columns If we prefer to get totals only on column level we can use: ```python df.pivot_table(index='foo', columns='bar', values='baz', margins=True).iloc[:, :-1] ``` Which results into total as a new column: | bar | A | B | C | All | | --- | --- | --- | --- | --- | | foo | | | | | | one | 1.0 | 2.0 | 3.0 | 2.0 | | two | 4.0 | 5.0 | 6.0 | 5.0 | ## Conclusion By default we can only control whether or not to have total columns by passing `margins`: > If margins=True, special All columns and rows will be added with partial group aggregates across the categories on the rows and columns. We can also change the name of the Total columns by: `margins_name`: > Name of the row / column that will contain the totals when margins is True. ## Resources - [Adding margins](https://pandas.pydata.org/docs/user%5Fguide/reshaping.html?ref=datascientyst.com#adding-margins) - [pandas.pivot\_table](https://pandas.pydata.org/docs/reference/api/pandas.pivot%5Ftable.html?ref=datascientyst.com) - [pandas.crosstab](https://pandas.pydata.org/docs/reference/api/pandas.crosstab.html?ref=datascientyst.com) ### TypeError: DataFrame.pivot() takes 1 positional argument but 4 were given - Pandas URL: https://datascientyst.com/typeerror-dataframe-pivot-takes-1-positional-argument-but-4-were-given-pandas/ Last updated: 2023-08-27T08:00:09.000Z In this tutorial, we'll take a closer look at the Pandas error, **TypeError: DataFrame.pivot() takes 1 positional argument but 4 were given - Pandas**. First, we'll create an example of how to produce it. Next, we'll explain the leading cause of the exception. And finally, we'll see how to fix it. ## Example Let's have a DataFrame like: ```python import pandas as pd df = pd.DataFrame({'foo': ['one', 'one', 'one', 'two', 'two', 'two'], 'bar': ['A', 'B', 'B', 'A', 'B', 'C'], 'baz': [1, 2, 3, 4, 5, 6], 'zoo': ['x', 'y', 'z', 'q', 'w', 't']}) ``` with data: | | foo | bar | baz | zoo | | - | --- | --- | --- | --- | | 0 | one | A | 1 | x | | 1 | one | B | 2 | y | | 2 | one | B | 3 | z | | 3 | two | A | 4 | q | | 4 | two | B | 5 | w | When running a code like the one below: ```python df.pivot('foo', 'bar', 'baz') ``` we get **error message: TypeError: DataFrame.pivot() takes 1 positional argument but 4 were given** ![TypeError: DataFrame.pivot() takes 1 positional argument but 4 were given - Pandas](https://datascientyst.com/content/images/2023/08/typeerror-dataframe-pivot-takes-1-positional-argument-but-4-were-given-pandas.png) ## Cause The reason for the error is that: > all arguments of DataFrame.pivot are keyword-only. Which means that we need to provide argument names for each argument passed to this method. You can read more on this link: [ENH: Keep positional arguments for pivot #51359](https://github.com/pandas-dev/pandas/issues/51359?ref=datascientyst.com) The code above was working in the past. Pandas community is doing efforts to make Pandas code more: - readable - explicit ## Solution To solve error - TypeError: DataFrame.pivot() takes 1 positional argument but 4 were given - we need to provide all arguments by name: ```python df.pivot_table(index='foo', columns='bar', values='baz') ``` This code will work fine. ## Conclusion We've explained **Pandas's TypeError: DataFrame.pivot() takes 1 positional argument but 4 were given error.** Then, we discussed the reason and shared a discussion on the topic. Lastly, we discussed how to resolve the error. ### Pandas pivot - ValueError: Index contains duplicate entries, cannot reshape URL: https://datascientyst.com/pandas-pivot-warning-about-repeated-entries-on-index/ Last updated: 2023-08-27T06:42:04.000Z In this article we will see how to solve **Pandas pivot error: "ValueError: Index contains duplicate entries, cannot reshape".** Let's see how to solve this error in different ways depending on the case. ## Setup Suppose we have a DataFrame like: ```python import pandas as pd df = pd.DataFrame({'foo': ['one', 'one', 'one', 'two', 'two', 'two'], 'bar': ['A', 'B', 'B', 'A', 'B', 'C'], 'baz': [1, 2, 3, 4, 5, 6], 'zoo': ['x', 'y', 'z', 'q', 'w', 't']}) ``` You can see data below: | | foo | bar | baz | zoo | | - | --- | --- | --- | --- | | 0 | one | A | 1 | x | | 1 | one | B | 2 | y | | 2 | one | B | 3 | z | | 3 | two | A | 4 | q | | 4 | two | B | 5 | w | | 5 | two | C | 6 | t | Note that there is a duplication in the row with index - 2 we have B instead of C. If we try to use method 'pivot' with duplicate entries like: ```python df.pivot(index='foo', columns='bar', values='baz') ``` we will get the error: `ValueError: Index contains duplicate entries, cannot reshape"` ![ValueError: Index contains duplicate entries, cannot reshape](https://datascientyst.com/content/images/2023/08/pandas-pivot-warning-about-repeated-entries-on-index.png) ## 1\. Use pivot\_table For tables with duplicate entries we need to use `pivot_table`: ```python df.pivot_table(index='foo', columns='bar', values='baz') ``` this will solve the error and produce correct result: | bar | A | B | C | | --- | --- | --- | --- | | foo | | | | | one | 1.0 | 2.5 | NaN | | two | 4.0 | 5.0 | 6.0 | ## 2\. Remove duplicates If you prefer to use the `pivot` method you need to drop duplicates from the DataFrame by: ```python df = df.drop_duplicates(['foo','bar']) df.pivot(index='foo', columns='bar', values='baz') ``` In this case the result is the same as using `pivot_table`: | bar | A | B | C | | --- | --- | --- | --- | | foo | | | | | one | 1.0 | 2.5 | NaN | | two | 4.0 | 5.0 | 6.0 | ## 3\. Aggregate You can also use a custom aggregation to mimic pivot behavior. Let's combine methods like: - `groupby` - `sum` to produce aggregate data as the method `pivot` without getting error: ```python df_agg = df.groupby(by=['foo', 'bar']).sum().reset_index() df_agg.pivot(index='foo', columns='bar', values='baz') ``` And again we get the same result: | bar | A | B | C | | --- | --- | --- | --- | | foo | | | | | one | 1.0 | 2.5 | NaN | | two | 4.0 | 5.0 | 6.0 | ## Conclusion In this post, we covered the most common solution for Pandas error on method pivot: > "ValueError: Index contains duplicate entries, cannot reshape". To solve Pandas errors you need to: - understand your data very well - know what the expected result should be. ## Resources - [DataFrame.pivot](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.pivot.html?ref=datascientyst.com) - [pandas.pivot\_table](https://pandas.pydata.org/docs/reference/api/pandas.pivot%5Ftable.html?ref=datascientyst.com) ### Football Prediction in Python: Barcelona vs Real Madrid URL: https://datascientyst.com/football-prediction-in-python-barcelona-vs-real-madrid/ Last updated: 2023-04-05T09:10:52.000Z In this post, we will Pandas and Python to collect football data and analyse it. We will try to predict probability for the outcome and the result of the fooball game between: Barcelona vs Real Madrid. Today is a great day for football fans - Barcelona vs Real Madrid game will be held tomorrow. Fans try to predict: - Who is better: Barca or Real? - Who won the most El Clasico? - Who will win the game? - Barcelona vs Real Madrid - Prediction, Odds and Betting Tips Can we use data science to find answers to those questions? Let's give it a try. **From this post you can learn**: - Scraping tables with Pandas - Pivot/unpivot tables - Pandas chaining - Fuzzy string matching in Python - Basic implementation in Python of: - Markov chain - ML model At the end there is a link to Python playbook in Kaggle. ## 1\. Collect stats Often things start with data collection. Nowadays it is much easier to collect data. Below you can find few ways to scrape football data with Python: ### Wikipedia - Historical data Wikipedia is a great source of information for El Clasico. We will use Pandas method `pd.read_html` to collect first table on the page: ```python import pandas as pd url = 'https://en.wikipedia.org/wiki/List_of_El_Cl%C3%A1sico_matches' df = pd.read_html(url)[0] df ``` data looks like: | | No. | Date | Matchweek | Home team | Away team | Score (FT/HT) | Goals (Home) | Goals (Away) | | - | --- | ---------------- | --------- | ----------- | ----------- | ------------- | ----------------------------------- | ------------------------------------------- | | 0 | 1 | 17 February 1929 | 2 | Barcelona | Real Madrid | 1–2 (0–1) | Parera (70) | Morera (10, 55) | | 1 | 2 | 9 May 1929 | 11 | Real Madrid | Barcelona | 0–1 (0–0) | NaN | Sastre (83) | | 2 | 3 | 26 January 1930 | 9 | Barcelona | Real Madrid | 1–4 (0–3) | Bestit (63) | Rubio (10, 37), F. López (17), Lazcano (71) | | 3 | 4 | 30 March 1930 | 18 | Real Madrid | Barcelona | 5–1 (3–0) | Rubio (5, 23), Lazcano (42, 68, 72) | Goiburu (84) | | 4 | 5 | 1 February 1931 | 9 | Real Madrid | Barcelona | 0–0 | NaN | NaN | ### Wikipedia - Current season To get the current results standings we can use wikipedia again with the following code: ```python import pandas as pd url = 'https://en.wikipedia.org/wiki/2022%E2%80%9323_La_Liga' df = pd.read_html(url)[6] df ``` You can find the table which we are scraping ![football-data-science-projects-real-vs-barcelona](https://datascientyst.com/content/images/2023/03/football-data-science-projects-real-vs-barcelona.webp) ### Google We can use google to search for the latests games and win probability by: [barcelona vs real madrid](https://www.google.com/search?q=real+vs+barcelona+who+will+win+&sxsrf=AJOqlzXJ3asA4dC20B5jRvUbgAXr3AbzvQ%3A1679216657537&ei=EdAWZJW9IImLxc8PyJqT-AY&ved=0ahUKEwjVif7C0ef9AhWJRfEDHUjNBG8Q4dUDCBA&uact=5&oq=real+vs+barcelona+who+will+win+&gs%5Flcp=Cgxnd3Mtd2l6LXNlcnAQAzIFCCEQoAEyBQghEKABMgUIIRCgATIFCCEQoAEyBQghEKABOgoIABBHENYEELADSgQIQRgAUMMCWMMCYJcWaAJwAXgAgAGcAYgBnAGSAQMwLjGYAQCgAQHIAQjAAQE&sclient=gws-wiz-serp&ref=datascientyst.com#sie=m;/g/11s3qt1995;2;/m/09gqx;dt;fp;1;;;). Which will give us: ![](https://datascientyst.com/content/images/2023/03/Football-Data-Science-Projects--Real-Vs-Barcelona.webp) I'm using browser extension to extract the table: [Table Capture](https://chrome.google.com/webstore/detail/table-capture/iebpjdmgckacbodjpijphcplhebcmeop?hl=en&ref=datascientyst.com) ## 2\. Transform Data Table below is stored in pivot or aggregated mode. In this step we will transform the data to a normal form. ![football-data-science-projects-real-vs-barcelona](https://datascientyst.com/content/images/2023/03/football-data-science-projects-real-vs-barcelona.webp) ### melt column names to values What we like to achieve is the following format: | | Home \\ Away | variable | value | corr | | - | --------------- | -------- | ----- | ---- | | 0 | Almería | ALM | — | VAL | | 1 | Athletic Bilbao | ALM | 4–0 | NaN | | 2 | Atlético Madrid | ALM | NaN | ATM | | 3 | Barcelona | ALM | 2–0 | NaN | | 4 | Cádiz | ALM | 1–1 | ELC | Pandas offer method `melt` which can turn columns into rows/values: ```python df_melt = pd.melt(df, id_vars=['Home \ Away']) df_melt ``` ## 3\. Match abbreviations and full names We have two different names of the teams: - full name - abbreviation We will use basic fuzzy string matching to map them - we will cover two ways for mapping: ### difflib Module `difflib` is very powerful and offers a quick way to match similar strings. In our case we can try to map abbreviations and full names by method `get_close_matches`: ```python teams = ['Real Madrid', 'Real Sociedad', 'Rayo Vallecano'] difflib.get_close_matches('RSO', teams, n=3, cutoff=0.2) ``` First we can collect all names and run test for similarity by: ```python full_names = df_melt['Home \ Away'].unique() abbr_names = df_melt['variable'].unique() for keyword in abbr_names: matches = difflib.get_close_matches(keyword, full_names, n=20, cutoff=0.2) print(keyword, matches) ``` Results are far for perfect: ``` ALM ['Atlético Madrid', 'Almería'] ATH ['Almería'] ATM ['Atlético Madrid', 'Almería'] BAR ['Almería'] CAD ['Cádiz', 'Almería'] CEL ['Elche', 'Cádiz'] ELC ['Elche', 'Cádiz'] ESP ['Elche', 'Sevilla'] GET ['Elche', 'Girona', 'Getafe'] ``` So the full code for Pandas to match the teams would be: ```python import difflib correct_values = {} words = full_names for keyword in abbr_names: similar = difflib.get_close_matches(keyword, words, n=3, cutoff=0.2) for x in similar: correct_values[x] = keyword df_melt["corr"] = df_melt["Home \ Away"].map(correct_values) ``` This method is very generic and produces results with errors. ### Regex to match name and abbreviation Let's try one more way to match abbreviations to possible names by using regex. This way is adjusted to the logic of abbreviating football teams. The code is: ```python import re def is_abbrev(abbrev, words): matches = [] for word in words: pattern = "(|.*\s)".join(abbrev.lower()) if re.match("^" + pattern, word.lower()) is not None: matches.append(word) return matches for keyword in abbr_names: matches = is_abbrev(keyword, full_names) print(keyword, matches) ``` and the results are much better: ``` ALM ['Almería'] ATH ['Athletic Bilbao'] ATM ['Atlético Madrid'] BAR ['Barcelona'] CAD [] CEL ['Celta Vigo'] ELC ['Elche'] ESP ['Espanyol'] GET ['Getafe'] GIR ['Girona'] MLL [] ``` Now we can export all games between Barcelona and Real Madrid. We can also find their latest games in Spain. ## 4\. Python: odds in sport This is a definition from wikipedia about odds: > In probability theory, odds provide a measure of the likelihood of a particular outcome. They are calculated as the ratio of the number of events that produce that outcome to the number that do not. You can find more about this in the Resource section. ### soccerapi Below you can find simple way to collect odds in Python by using library -`soccerapi` \- `pip install soccerapi`: ```python from soccerapi.api import Api888Sport api = Api888Sport() url = 'https://www.888sport.com/#/filter/football/spain/' odds = api.odds(url) pd.DataFrame(odds) ``` This will collect all odds for Spain: | | time | home\_team | away\_team | full\_time\_result | under\_over | both\_teams\_to\_score | double\_chance | | - | -------------------- | --------------- | ----------- | --------------------------------- | ---------------------------- | ------------------------- | ------------------------------------ | | 0 | 2023-04-04T19:00:00Z | Athletic Bilbao | Osasuna | {'1': 1460, 'X': 4000, '2': 7500} | {'O2.5': 2140, 'U2.5': 1730} | {'yes': 2250, 'no': 1610} | {'1X': 1110, '12': 1260, '2X': 2500} | | 1 | 2023-04-05T19:00:00Z | FC Barcelona | Real Madrid | {'1': 2250, 'X': 3400, '2': 3100} | {'O2.5': 1830, 'U2.5': 2000} | {'yes': 1650, 'no': 2180} | {'1X': 1380, '12': 1330, '2X': 1610} | | 2 | 2023-04-07T16:30:00Z | Lugo | Tenerife | {'1': 4100, 'X': 2750, '2': 2020} | {'O2.5': 1630, 'U2.5': 2200} | {'yes': 2430, 'no': 1520} | {'1X': 1740, '12': 1410, '2X': 1220} | | 3 | 2023-04-07T19:00:00Z | Villarreal B | Málaga | {'1': 1930, 'X': 3250, '2': 3700} | {'O2.5': 2200, 'U2.5': 1630} | {'yes': 1980, 'no': 1780} | {'1X': 1260, '12': 1320, '2X': 1810} | | 4 | 2023-04-07T19:00:00Z | Sevilla | Celta Vigo | {'1': 2150, 'X': 3300, '2': 3400} | {'O2.5': 2250, 'U2.5': 1640} | {'yes': 1980, 'no': 1780} | {'1X': 1330, '12': 1340, '2X': 1670} | We can find the odds for: FC Barcelona and Real Madrid: `{'1': 2250, 'X': 3400, '2': 3100}`. ### sports-betting Another useful Python package for betting is: `sports-betting`. It can be installed by: `pip install sports-betting` This package require Python 3.9+ and offer simple API: ```python from sportsbet.datasets import SoccerDataLoader dataloader = SoccerDataLoader(param_grid={'league': ['Italy'], 'year': [2020]}) X_train, Y_train, O_train = dataloader.extract_train_data(odds_type='market_maximum', drop_na_thres=1.0) X_fix, Y_fix, O_fix = dataloader.extract_fixtures_data() ``` ### Calculate odds We can follow the next steps in order to calculate the odds in a custom code: - Barcelona has a home advantage. - Barcelona has (out of their last 10 games): - won 7 - drawn 1 - lost 2 - Real Madrid: - won 6 - drawn 2 - lost 2 To calculate the odds for this match, we can use simple Python code: ```python barca_win_rate = 7 / 10 real_win_rate = 6 / 10 draw_rate = 1 - barca_win_rate - real_win_rate barca_odds = 1 / barca_win_rate real_odds = 1 / real_win_rate draw_odds = 1 / draw_rate print("Barcelona odds:", barca_odds) print("Real Madrid odds:", real_odds) print("Draw odds:", draw_odds) ``` which give us: ``` Barcelona odds: 1.4285714285714286 Real Madrid odds: 1.6666666666666667 Draw odds: -3.333333333333334 ``` ## 5\. Python Markov Chain Finally we can use Markov Chains to calculate probability for win, draw and lose. ### Collect data We will collect all previous el clasico games by: ```python import pandas as pd url = 'http://eurorivals.net/head-to-head/barcelona-vs-real-madrid' df_el_cl = pd.read_html(url)[1] df_el_cl.head() ``` result: | | Unnamed: 0 | Unnamed: 1 | Unnamed: 2 | Unnamed: 3 | Unnamed: 4 | Unnamed: 5 | | - | --------------- | ---------- | ---------- | ----------- | ---------- | ----------- | | 0 | 2023 19 March | NaN | Primera | Barcelona | 2 - 1 | Real Madrid | | 1 | 2023 2 March | NaN | Cup | Real Madrid | 0 - 1 | Barcelona | | 2 | 2022 16 October | NaN | Primera | Real Madrid | 3 - 1 | Barcelona | | 3 | 2022 20 March | NaN | Primera | Real Madrid | 0 - 4 | Barcelona | | 4 | 2021 24 October | NaN | Primera | Barcelona | 1 - 2 | Real Madrid | ### Data cleaning with Panda chaining In this step we will demonstrate how to use Pandas chaining for data cleaning and preprocessing: ```python import numpy as np df_trans = ( df .set_axis(['date', 'flag', 'tournament', 'home', 'score', 'away'], axis=1) # rename columns .drop('flag', axis=1) # drop columns .assign(score_home=lambda x: x.score.str.split('-', expand=True)[0].astype(int)) # split and add new column .assign(score_away=lambda x: x.score.str.split('-', expand=True)[1].astype(int)) # split and add new column .assign(date_n=lambda x: pd.to_datetime(x.date)) # convert to datetime .assign(goal_diff=lambda x: x.score_home - x.score_away) # compare two columns .assign(result= lambda x: x['goal_diff'].apply(lambda y: "L" if y < 0 else ("W" if y > 0 else "D"))) # conditional in chaining .set_index('date_n') # set new index .sort_index() # sort by index .loc[:,['home', 'away', 'score_home', 'score_away', 'result']] # return column from chaining ) ``` After the preprocessing we have this data: | | home | away | score\_home | score\_away | result | | ---------- | ----------- | ----------- | ----------- | ----------- | ------ | | date\_n | | | | | | | 2004-04-24 | Real Madrid | Barcelona | 1 | 2 | L | | 2004-11-19 | Barcelona | Real Madrid | 3 | 0 | W | | 2005-04-09 | Real Madrid | Barcelona | 4 | 2 | W | | 2005-11-18 | Real Madrid | Barcelona | 0 | 3 | L | | 2006-03-31 | Barcelona | Real Madrid | 1 | 1 | D | | 2006-10-22 | Real Madrid | Barcelona | 2 | 0 | W | | 2007-03-10 | Barcelona | Real Madrid | 3 | 3 | D | | 2007-12-23 | Barcelona | Real Madrid | 0 | 1 | L | | 2008-05-07 | Real Madrid | Barcelona | 4 | 1 | W | | 2008-12-13 | Barcelona | Real Madrid | 2 | 0 | W | | 2009-05-02 | Real Madrid | Barcelona | 2 | 6 | L | | 2009-11-29 | Barcelona | Real Madrid | 1 | 0 | W | | 2010-04-10 | Real Madrid | Barcelona | 0 | 2 | L | | 2010-11-29 | Barcelona | Real Madrid | 5 | 0 | W | | 2011-04-16 | Real Madrid | Barcelona | 1 | 1 | D | | 2011-04-20 | Barcelona | Real Madrid | 0 | 0 | D | | 2011-04-27 | Real Madrid | Barcelona | 0 | 2 | L | | 2011-05-03 | Barcelona | Real Madrid | 1 | 1 | D | | 2011-12-10 | Real Madrid | Barcelona | 1 | 3 | L | | 2012-01-18 | Real Madrid | Barcelona | 1 | 2 | L | | 2012-01-25 | Barcelona | Real Madrid | 1 | 2 | L | | 2012-04-21 | Barcelona | Real Madrid | 1 | 2 | L | | 2012-10-07 | Barcelona | Real Madrid | 2 | 2 | D | | 2013-01-30 | Real Madrid | Barcelona | 1 | 1 | D | | 2013-02-26 | Barcelona | Real Madrid | 1 | 3 | L | | 2013-03-02 | Real Madrid | Barcelona | 2 | 1 | W | | 2013-10-26 | Barcelona | Real Madrid | 2 | 1 | W | | 2014-03-23 | Real Madrid | Barcelona | 3 | 4 | L | | 2014-04-16 | Barcelona | Real Madrid | 1 | 2 | L | | 2014-10-25 | Real Madrid | Barcelona | 3 | 1 | W | | 2015-03-22 | Barcelona | Real Madrid | 2 | 1 | W | | 2015-11-21 | Real Madrid | Barcelona | 0 | 4 | L | | 2016-04-02 | Barcelona | Real Madrid | 1 | 2 | L | | 2016-12-03 | Barcelona | Real Madrid | 1 | 1 | D | | 2017-04-23 | Real Madrid | Barcelona | 2 | 3 | L | | 2017-12-23 | Real Madrid | Barcelona | 0 | 3 | L | | 2018-05-06 | Barcelona | Real Madrid | 2 | 2 | D | | 2018-10-28 | Barcelona | Real Madrid | 5 | 1 | W | | 2019-02-06 | Barcelona | Real Madrid | 1 | 1 | D | | 2019-02-27 | Real Madrid | Barcelona | 0 | 3 | L | | 2019-03-02 | Real Madrid | Barcelona | 0 | 1 | L | | 2019-12-18 | Barcelona | Real Madrid | 0 | 0 | D | | 2020-03-01 | Real Madrid | Barcelona | 2 | 0 | W | | 2020-10-24 | Barcelona | Real Madrid | 1 | 3 | L | | 2021-04-10 | Real Madrid | Barcelona | 2 | 1 | W | | 2021-10-24 | Barcelona | Real Madrid | 1 | 2 | L | | 2022-03-20 | Real Madrid | Barcelona | 0 | 4 | L | | 2022-10-16 | Real Madrid | Barcelona | 3 | 1 | W | | 2023-03-02 | Real Madrid | Barcelona | 0 | 1 | L | | 2023-03-19 | Barcelona | Real Madrid | 2 | 1 | W | ### Markov chains probability We will use package `mchmm` which can be installed by: `pip install mchmm` In order to find the transition matrix and plot graph of probability changes: ```python import mchmm as mc a = mc.MarkovChain().from_data(df_trans['result']) ``` So we get probability matrix by: ```python a.observed_p_matrix ``` results into: array(\[\[0.18181818, 0.54545455, 0.27272727\], \[0.26086957, 0.34782609, 0.39130435\], \[0.2 , 0.53333333, 0.26666667\]\]) or as a table: | | 0 | 1 | 2 | | - | -------- | -------- | -------- | | 0 | 0.181818 | 0.545455 | 0.272727 | | 1 | 0.260870 | 0.347826 | 0.391304 | | 2 | 0.200000 | 0.533333 | 0.266667 | And the following transition graph by: ```python graph = a.graph_make( format="png", graph_attr=[("rankdir", "LR")], node_attr=[("fontname", "Roboto bold"), ("fontsize", "20")], edge_attr=[("fontname", "Iosevka"), ("fontsize", "12")] ) graph.render() ``` graph: ![markov_chain_graph_football_prediction](https://datascientyst.com/content/images/2023/04/markov_chain_graph_football_prediction.png) So the draw state seems to be less favorable. After winning, often there is a loss. Note: We don't take into account home and away teams. ## 6\. Predict score with simple ML model Finally we can try to predict the score based on the scores so far. We can use simple ML model with: - inputs - home and away teams - outputs - home and away scores The code is: ```python from sklearn.linear_model import LinearRegression mapping = {'Barcelona':1, 'Real Madrid':2} df_trans = df_trans.replace(mapping).head() TRAIN_INPUT = df_trans[['home', 'away']].values TRAIN_OUTPUT = df_trans[['score_home', 'score_away']].values X_TEST = [[2, 1]] outcome = predictor.predict(X=X_TEST) coefficients = predictor.coef_ print(outcome) print('Outcome : Coefficients : {}'.format(coefficients)) ``` which give us: \[\[1.66666667 2.33333333\]\] Outcome : Coefficients : \[\[-0.16666667 0.16666667\] \[ 0.91666667 -0.91666667\]\] Note: this is a pretty naive way to try to predict score in Football. ## 7\. Resources Helpful resources for aspiring data scientists who would like to research football by data science deeper: - [Notebook](https://www.kaggle.com/softhints/football-prediction-barcelona-vs-real-madrid?ref=datascientyst.com) - [How to Melt Pandas DataFrame](https://datascientyst.com/use-melt-pandas-dataframe-pd-melt-examples/) - [difflib — Helpers for computing deltas](https://docs.python.org/3/library/difflib.html?ref=datascientyst.com) - Data Source - [Barca, Madrid and Visualisation](https://www.kaggle.com/code/adikeshri/barca-madrid-and-visualisation/comments?ref=datascientyst.com) - [El Clásico](https://en.wikipedia.org/wiki/El%5FCl%C3%A1sico?ref=datascientyst.com) - [List of El Clásico matches](https://en.wikipedia.org/wiki/List%5Fof%5FEl%5FCl%C3%A1sico%5Fmatches?ref=datascientyst.com) - Odds - [Odds ratio](https://en.wikipedia.org/wiki/Odds%5Fratio?ref=datascientyst.com) - [Statistical association football predictions](https://en.wikipedia.org/wiki/Statistical%5Fassociation%5Ffootball%5Fpredictions?ref=datascientyst.com) - [Odds](https://en.wikipedia.org/wiki/Odds?ref=datascientyst.com) - [Odds != Probability](https://towardsdatascience.com/odds-probability-c9cf80405027?ref=datascientyst.com) - Python packages - [soccerapi](https://pypi.org/project/soccerapi/?ref=datascientyst.com) \- wrapper build on top of some bookmakers (888sport, bet365 and Unibet) in order to get data about soccer (aka football) odds using python commands - [sports-betting](https://pypi.org/project/sports-betting/?ref=datascientyst.com) \- collection of tools that makes it easy to create machine learning models for sports betting - [mchmm](https://pypi.org/project/mchmm/?ref=datascientyst.com) \- Markov chains and Hidden Markov models ### How to solve: HTTPError: HTTP Error 403: Forbidden in Pandas URL: https://datascientyst.com/how-to-solve-httperror-http-error-403-forbidden-in-pandas/ Last updated: 2023-04-03T14:23:53.000Z In this post you can find how to solve Pandas and Python error: > HTTPError: HTTP Error 403: Forbidden ## HTTPError: HTTP Error 403: Forbidden This error happens when we try to scrape tables with Pandas by using `read_html` method. For example: ```python import pandas as pd url_cur = 'https://tradingeconomics.com/currencies' pd.read_html(url_cur)[0] ``` This results into error - **HTTPError: HTTP Error 403: Forbidden**. To solve this error we can simulate browser and user agent in Pandas by passing headers. ## Solution Below you can find how to fix the error: ```python import requests import pandas as pd url_cur = 'https://tradingeconomics.com/currencies' header = { "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/50.0.2661.75 Safari/537.36", "X-Requested-With": "XMLHttpRequest" } r = requests.get(url_cur, headers=header) pd.read_html(r.text)[0] ``` result: | | Unnamed: 0 | Major | Price | Day | % | Weekly | Monthly | YoY | Date | | - | ---------- | ------ | --------- | ------ | ------- | ------- | ------- | -------- | ------ | | 0 | NaN | EURUSD | 1.08909 | 0.0052 | 0.48% | 0.88% | 1.99% | \-0.72% | Apr/03 | | 1 | NaN | GBPUSD | 1.24050 | 0.0072 | 0.58% | 0.99% | 3.19% | \-5.40% | Apr/03 | | 2 | NaN | AUDUSD | 0.67726 | 0.0088 | 1.31% | 1.86% | 0.68% | \-10.20% | Apr/03 | | 3 | NaN | NZDUSD | 0.62818 | 0.0025 | 0.40% | 1.42% | 1.42% | \-9.61% | Apr/03 | | 4 | NaN | USDJPY | 132.53900 | 0.2510 | \-0.19% | 0.74% | \-2.48% | 7.95% | Apr/03 | | 5 | NaN | USDCNY | 6.88090 | 0.0068 | 0.10% | \-0.01% | \-0.99% | 7.97% | Apr/03 | | 6 | NaN | USDCHF | 0.91320 | 0.0016 | \-0.17% | \-0.27% | \-1.88% | \-1.41% | Apr/03 | We solve the error by: - using requests module - download the page by using headers - parse downloaded data with Pandas ## Pandas authorization by user and password Sometimes you may need to log by using user and password. The example below shows how to use `requests` library to perform such request: ```python import requests import pandas as pd url = 'https://example.com' username = 'your_username' password = 'your_password' response = requests.get(url, auth=(username, password)) if response.status_code == 200: df = pd.read_html(response.content)[0] print(df.head()) else: print(f'Request failed') ``` ### How to Round Time to the Nearest Quarter or Hour in Pandas? URL: https://datascientyst.com/how-to-round-time-to-the-nearest-quarter-or-hour-in-pandas/ Last updated: 2023-04-02T20:38:19.000Z To round a **datetime column to the nearest quarter, minute or hour in Pandas**, we can use the method: `dt.round()`. ![](https://datascientyst.com/content/images/2023/04/round-time-to-the-nearest-quarter-or-hour-in-pandas.png) ## round datetime column to nearest hour Below you can find an example of rounding to the closest hour in Pandas and Python. We use method `dt.round()` with parameter `H`: ```python import pandas as pd dates = ['2023-03-25 11:37:00', '2023-03-25 09:18:00', '2023-03-25 15:23:00'] df = pd.DataFrame({'date': dates}) df['date'] = pd.to_datetime(df['date']) df['date'].dt.round('H') ``` the result is rounded times to the nearest hour: ``` 0 2023-03-25 12:00:00 1 2023-03-25 09:00:00 2 2023-03-25 15:00:00 Name: date, dtype: datetime64[ns] ``` ## round datetime to nearest quarter or minutes We can round to the nearest quarter in Pandas by the same method: `.dt.round('15min')` specifying the interval in minutes. Example of rounding down in Pandas to N minutes: ```python import pandas as pd dates = ['2023-03-25 11:37:00', '2023-03-25 09:18:00', '2023-03-25 15:23:00'] df = pd.DataFrame({'date': dates}) df['date'] = pd.to_datetime(df['date']) df['date'].dt.round('15min') ``` result of rounding to quarter is: ``` 0 2023-03-25 11:30:00 1 2023-03-25 09:15:00 2 2023-03-25 15:30:00 Name: date, dtype: datetime64[ns] ``` In the next section you can find a link to all possible frequency values. ## datetime - round vs floor Finally let's see what is the difference between Pandas methods: - round - floor ```python df['rounded_date'] = df['date'].dt.round('H') df['floored_date'] = df['date'].dt.floor('H') ``` You can find the result below: | | date | rounded\_date | floored\_date | | - | ------------------- | ------------------- | ------------------- | | 0 | 2023-03-25 11:37:00 | 2023-03-25 12:00:00 | 2023-03-25 11:00:00 | | 1 | 2023-03-25 09:18:00 | 2023-03-25 09:00:00 | 2023-03-25 09:00:00 | | 2 | 2023-03-25 15:23:00 | 2023-03-25 15:00:00 | 2023-03-25 15:00:00 | So difference is in the first row where: - `12:00:00` \- round the datetime column to the nearest hour - `11:00:00` \- floor the datetime column to the nearest hour ## Resources - [pandas.Series.dt.round](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.round.html?ref=datascientyst.com) - [pandas.to\_datetime](https://pandas.pydata.org/docs/reference/api/pandas.to%5Fdatetime.html?highlight=to%5Fdatetime&ref=datascientyst.com) - [Time series / date functionality](https://pandas.pydata.org/docs/user%5Fguide/timeseries.html?ref=datascientyst.com) - [Offset/frequency aliases](https://pandas.pydata.org/docs/user%5Fguide/timeseries.html?ref=datascientyst.com#offset-aliases) \- for a list of possible freq values - [pandas.Series.dt.floor](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.floor.html?ref=datascientyst.com) ### How to Sort by Multiple Columns Ascending and Descending in Pandas? URL: https://datascientyst.com/how-to-sort-by-multiple-columns-ascending-and-descending-in-pandas/ Last updated: 2023-04-02T20:21:19.000Z To sort by multiple columns ascending and descending in Pandas we can use syntax like: ```python df.sort_values(by=['name', 'salary'], ascending=[True, False]) ``` ![](https://datascientyst.com/content/images/2023/04/sort-by-multiple-columns-ascending-and-descending-in-pandas.webp) Let's cover two examples to explain sorting on multiple columns in more detail. ## Sort a DataFrame by two or more columns To sort Pandas DataFrame by two and more columns we can use parameter `by`: `df.sort_values(by=['name', 'age'], ascending=True)` Example of sorting DataFrame by two columns: ```python import pandas as pd df = pd.DataFrame({ 'name': ['Alice', 'Bob', 'Charlie', 'Alice', 'Bob'], 'age': [25, 30, 35, 25, 40], 'salary': [50000, 70000, 60000, 55000, 80000] }) df.sort_values(by=['name', 'age'], ascending=True) ``` The code above sort the DataFrame by 'name' and 'age' columns in ascending order: | | name | age | salary | | - | ------- | --- | ------ | | 0 | Alice | 25 | 50000 | | 3 | Alice | 25 | 55000 | | 1 | Bob | 30 | 70000 | | 4 | Bob | 40 | 80000 | | 2 | Charlie | 35 | 60000 | DataFrame before sorting is: | | name | age | salary | | - | ------- | --- | ------ | | 0 | Alice | 25 | 50000 | | 1 | Bob | 30 | 70000 | | 2 | Charlie | 35 | 60000 | | 3 | Alice | 25 | 55000 | | 4 | Bob | 40 | 80000 | We can specify the sort order for each column using the ascending parameter as boolean or a list: - `ascending=True` \- the columns will be sorted in ascending order - `ascending=False` \- sorted in descending order. ## sort by multiple columns one ascending and descending We can use Pandas method `sort_values()` to sort by multiple columns in different order: ascending and descending. Parameter `ascending` can take a list of values: `ascending=[True, False]`. Full example of sort both ascending and descending in Pandas: ```python import pandas as pd df = pd.DataFrame({ 'name': ['Alice', 'Bob', 'Charlie', 'Alice', 'Bob'], 'age': [25, 30, 35, 25, 40], 'salary': [50000, 70000, 60000, 55000, 80000] }) df.sort_values(by=['name', 'salary'], ascending=[True, False]) ``` The result of the sorted DataFrame is: | | name | age | salary | | - | ------- | --- | ------ | | 3 | Alice | 25 | 55000 | | 0 | Alice | 25 | 50000 | | 4 | Bob | 40 | 80000 | | 1 | Bob | 30 | 70000 | | 2 | Charlie | 35 | 60000 | ## Resources - [pandas.DataFrame.sort\_values](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.sort%5Fvalues.html?ref=datascientyst.com) - [pandas.DataFrame.sort\_index](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.sort%5Findex.html?ref=datascientyst.com) ### Convert Pivot Table to Regular Data Frame in Pandas URL: https://datascientyst.com/convert-pivot-table-to-regular-data-frame-in-pandas/ Last updated: 2023-04-01T06:43:52.000Z In this post, we will see how to convert a Pandas pivot table to a regular DataFrame. To convert pivot table to DataFrame we can use: **(1) the `reset_index()` method** ```python df_p.set_axis(df_p.columns.tolist(), axis=1).reset_index() ``` **(2) to\_records() + pd.DataFrame()** ```python pd.DataFrame(df_p.to_records()) ``` Let's cover both ways in detail in the next sections. ![](https://datascientyst.com/content/images/2023/04/convert-pivot-table-to-regular-data-frame-in-pandas.webp) ## Setup First, let's create the DataFrame with 4 columns: ```python import pandas as pd import numpy as np data = {'date': np.random.choice([202303, 202304], size=100), 'code': np.random.choice([*'ABC'], size=100), 'type': np.random.choice([*'RB'], size=100), 'val': np.arange(100)} df = pd.DataFrame(data) df ``` First 5 rows of this DataFrame are: | | date | code | type | val | | - | ------ | ---- | ---- | --- | | 0 | 202304 | B | B | 0 | | 1 | 202304 | A | B | 1 | | 2 | 202304 | B | B | 2 | | 3 | 202303 | C | B | 3 | | 4 | 202304 | A | B | 4 | We can create pivot table from above data by: ```python df_p = df.pivot_table(index=['date','code'], columns='type', values='val', aggfunc="count") df_p ``` result: | | type | B | R | | ------ | ---- | -- | - | | date | code | | | | 202303 | A | 9 | 5 | | B | 9 | 10 | | | C | 11 | 12 | | | 202304 | A | 11 | 6 | | B | 10 | 4 | | | C | 6 | 7 | | Let's see how to convert the pivot table back to normal DataFrame. Essentially this means to remove the MultiIndex or flatten the DataFrame. ## reset\_index() To convert pivot table to a normal DataFrame in Pandas, we can combine: - `reset_index()` method - `set_axis()` We can flatten the pivot table by removing the MultiIndex: ```python df_p.set_axis(df_p.columns.tolist(), axis=1).reset_index() ``` The result is: | | date | code | B | R | | - | ------ | ---- | -- | -- | | 0 | 202303 | A | 9 | 5 | | 1 | 202303 | B | 9 | 10 | | 2 | 202303 | C | 11 | 12 | | 3 | 202304 | A | 11 | 6 | | 4 | 202304 | B | 10 | 4 | | 5 | 202304 | C | 6 | 7 | ## to\_records() Another way to convert pivot tables in Pandas is by: - extracting data with `to_records()` - create new DataFrame ```python pd.DataFrame(df_p.to_records()) ``` We get the same result: | | date | code | B | R | | - | ------ | ---- | -- | -- | | 0 | 202303 | A | 9 | 5 | | 1 | 202303 | B | 9 | 10 | | 2 | 202303 | C | 11 | 12 | | 3 | 202304 | A | 11 | 6 | | 4 | 202304 | B | 10 | 4 | | 5 | 202304 | C | 6 | 7 | ## Summary We saw how to convert pivot tables to normal DataFrame in Pandas. This is useful when we need to work with pivot tables as regular DataFrame with MultiIndex. If you like to unpivot tables you can check the resources below. ## Resources - [reset\_index](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.reset%5Findex.html?ref=datascientyst.com) - [to\_records](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to%5Frecords.html?highlight=to%5Frecords&ref=datascientyst.com) - [How to Melt Pandas DataFrame - pd.melt in Examples](https://datascientyst.com/use-melt-pandas-dataframe-pd-melt-examples/) - [Opposite of Melt in Python and Pandas](https://datascientyst.com/opposite-of-melt-python-pandas/) - [How To Create a Pivot Table in Pandas?](https://datascientyst.com/how-to-create-a-pivot-table-in-pandas/) - [How to Flatten a MultiIndex in Pandas](https://datascientyst.com/flatten-multiindex-in-pandas/) ### Error: need to escape, but no escapechar set - Pandas URL: https://datascientyst.com/error-need-to-escape-but-no-escapechar-set-pandas/ Last updated: 2023-03-18T22:26:20.000Z In this tutorial, we'll see how to solve a common Pandas and Python error – "Error: need to escape, but no escapechar set". We get this error from Pandas when we try to save DataFrame as a CSV file. Let's see several examples of how to reproduce and solve this error. ## Pandas - Error: need to escape, but no escapechar set When we try to use `to_csv` we may get error in Pandas: > Error: need to escape, but no escapechar set The example below demonstrate the error: ```python import pandas import csv data = {'A': ['1.0627', '0625', '"LTE,eHRPD"'], 'B': ['0.3', '', '"ZTE\A"'], } df = pandas.DataFrame(data) display(df) print(df.to_csv(quoting=csv.QUOTE_NONE)) ``` The error is caused by using `csv.QUOTE_NONE`. Let's check two different solutions to this error. ## Pandas to\_csv escapechar First way to solve the error is by using `escapechar='\\'`: ```python print(df.to_csv(quoting=csv.QUOTE_NONE, quotechar='',escapechar='\\')) ``` This will produce CSV file as: ``` ,A,B 0,1.0627,0.3 1,0625, 2,"LTE\,eHRPD","ZTE\\A" ``` ## Solution - QUOTE\_ALL and QUOTE\_MINIMAL Alternatively we can solve the error by setting parameters: - `quoting=csv.QUOTE_MINIMAL` \- to only quote fields that contain special characters. - `quoting=csv.QUOTE_ALL` \- all fields will be quoted ### QUOTE\_MINIMAL We can see the output of: ```python df.to_csv(quoting=csv.QUOTE_MINIMAL) ``` the will be quoting the values in row 2: ``` ,A,B 0,1.0627,0.3 1,0625, 2,"""LTE,eHRPD""","""ZTE\A""" ``` ### QUOTE\_ALL We can see the output of - `QUOTE_ALL`: ```python df.to_csv(quoting=csv.QUOTE_ALL) ``` then all values will be quoted: ``` "","A","B" "0","1.0627","0.3" "1","0625","" "2","""LTE,eHRPD""","""ZTE\A""" ``` ![](https://datascientyst.com/content/images/2023/03/error-need-to-escape-but-no-escapechar-set-pandas.webp) ## Python - Error: need to escape, but no escapechar set The code below shows how to reproduce Python error - "Error: need to escape, but no escapechar set" ### QUOTE\_MINIMAL ```python import csv data = [["John", "Doe", "john@example.com"], ["Jane", "Doe", "jane@example.com", "Hello, \"world\"!"]] with open("output.csv", "w", newline="") as f: writer = csv.writer(f, quoting=csv.QUOTE_NONE) for row in data: writer.writerow(row) ``` result: ``` Error: need to escape, but no escapechar set ``` So the error is solved by: ```python writer = csv.writer(f, quoting=csv.QUOTE_MINIMAL) ``` ### Filter quoted values As alternative solution we can filter the values by list comprehension: ```python import csv data = [["John", "", "Doe"], ["Jane", " ", "Doe"]] with open("output.csv", "w", newline="") as f: writer = csv.writer(f) for row in data: row = [s if s else " " for s in row] # filter problematic strings writer.writerow(row) ``` ## Summary To summarize, in this article, we've seen how to solve Python and Pandas error: > Error: need to escape, but no escapechar set And finally you can find useful resources which will help you to solve Pandas `to_csv` output quoting issues. ## Resources - [\_csv.Error: need to escape, but no escapechar set](https://github.com/pandas-dev/pandas/issues/16298?ref=datascientyst.com) - [to\_csv](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to%5Fcsv.html?ref=datascientyst.com) - [Python CSV File Reading and Writing](https://docs.python.org/3/library/csv.html?ref=datascientyst.com) ### How to Read Data from Text File Into Pandas? URL: https://datascientyst.com/load-data-from-text-file-into-pandas/ Last updated: 2023-03-18T10:42:53.000Z The following step-by-step example shows how to load data from a text file into Pandas. We can use: - `read_csv()` function - it handles various delimiters, including commas, tabs, and spaces - `pd.read_fwf()` - read fixed-width formatted lines into DataFrame Let's cover both cases into examples: ## read\_csv - delimited file To read a text into Pandas DataFrame we can use method `read_csv()` and provide the separator: ```python import pandas as pd df = pd.read_csv('data.txt', sep=',') ``` Where `sep` argument specifies the separator. Separator can be continuous - `'\s+'`. Other useful parameters are: - `header=None` \- does the file contain headers - `names=["a", "b", "c"]` \- the column names - `skiprows=[0,1]` \- skip rows - `index_col=True` \- use index from the file ## read\_fwf - fixed-width file To read data from a fixed-width file in Pandas we can use [read\_fwf](https://pandas.pydata.org/docs/reference/api/pandas.read%5Ffwf.html?ref=datascientyst.com). Suppose we have a file `'data.txt'` like: ``` John 35 123 A Jane D 28 45 E Bob 42 678 D ``` We can see that columns are aligned by position rather than separated by delimiters. Since the columns are separated by fixed widths: - first column - 7 chars - (0, 7) - second - 2 chars - (7, 9) we can't use `read_csv()` with a separator. Instead we will: - specify the column widths - read the fixed-width file into a DataFrame ```python import pandas as pd colspecs = [(0, 7), (7, 9), (13, 17), (17, 18)] df = pd.read_fwf('data.txt', colspecs=colspecs, header=None, names=['name', 'age', 'score', 'class']) df ``` The result is: | | name | age | score | class | | - | ------ | --- | ----- | ----- | | 0 | John | 35 | 123 | A | | 1 | Jane D | 28 | 45 | E | | 2 | Bob | 42 | 678 | D | ![](https://datascientyst.com/content/images/2023/03/load-data-from-text-file-into-pandas.webp) ## Pandas read text file line by line To read a text file line by line into a pandas DataFrame we can: - create an empty DataFrame - create an iterator to read the file line by line - iterate over the iterator and append each line to the DataFrame - reset the index of the DataFrame ```python import pandas as pd df = pd.DataFrame() iterator = pd.read_csv('data.txt', header=None, iterator=True, chunksize=1) for chunk in iterator: df = df.append(chunk) df = df.reset_index(drop=True) ``` ## Pandas read text file with pattern As an alternative we can use list comprehension to read files and filter it. Let's work with the following file: ``` John 35 123 A Pattern Jane D 28 45 E Bob 42 678 D End of pattern ``` We can find the numbers of the start and end lines by matching pattern: ```python a=[] with open('data.txt',"r") as r: a=r.readlines() a=[x.replace("\n","") for x in a] start = a.index("Pattern") +1 end = a.index("End of pattern") start, end ``` After that we can read the file with `read_fwf` or `read_csv` and filter the lines: ```python import pandas as pd df = pd.read_fwf('data.txt', colspecs=colspecs, header=None, names=['name', 'age', 'score', 'class']) df = df[start:end] ``` Which give us: | | name | age | score | class | | - | ------ | --- | ----- | ----- | | 2 | Jane D | 28 | 45 | E | | 3 | Bob | 42 | 678 | D | ## Summary We've seen three different ways of reading and loading text file into Pandas DataFrame. We covered how to read delimited or fixed-length files with Pandas. We also saw how to read text files line by line and how to filter csv or text file by pattern. ### How to Extract Dictionary Value from Column in Pandas URL: https://datascientyst.com/extract-dictionary-value-from-column-in-pandas/ Last updated: 2023-03-18T09:44:49.000Z In this short guide, I'll show you how to **extract or explode a dictionary value from a column in a Pandas DataFrame**. You can use: - list or dict comprehension to extract dictionary values - the `apply()` function along with a lambda function to extract the value from each dictionary ![](https://datascientyst.com/content/images/2023/03/extract-dictionary-value-from-column-in-dataframe.webp) ## Setup For example, suppose you have a DataFrame with a column containing dictionaries: ```python import pandas as pd data = {'data': [{'a': 1, 'b': 2}, {'a': 3, 'b': 4}]} df = pd.DataFrame(data) ``` | | data | | - | ---------------- | | 0 | {'a': 1, 'b': 2} | | 1 | {'a': 3, 'b': 4} | ## Extract key from dict column To extract values for a particular key from the dictionary in each row, we can use the following code: ```python df['key'] = df['data'].apply(lambda x: x[key]) ``` which give us a Series of all values matching this key: ``` 0 1 1 3 Name: data, dtype: int64 ``` Finally we create a new column from the extracted data. As we mentioned at the start we can use list comprehension to extract all values to list: ```python [d.get('a') for d in df.data] ``` which give us: ``` [1, 3] ``` ## Split/Explode a column of dictionaries To split or explode a column of dictionaries to separate columns we can use: `.apply(pd.Series)`: ```python df['data'].apply(pd.Series) ``` this give us new DataFrame with columns from the exploded dictionaries: | | a | b | | - | - | - | | 0 | 1 | 2 | | 1 | 3 | 4 | A faster way to achieve similar behavior is by using `pd.json_normalize(df['data'])`: ```python pd.json_normalize(df['data']) ``` For more information and examples you can check: - [How to Normalize JSON or Dict to New Columns in Pandas ](https://datascientyst.com/normalize-json-dict-new-columns-pandas/) ## Extract all values/keys from dict column To extract all keys and values from a column which contains dictionary data we can list all keys by: `list(x.keys())`. Below are several examples: ```python df['keys'] = df['data'].apply(lambda x: list(x.keys())) df['values'] = df['data'].apply(lambda x: list(x.values())) ``` which create new column only with the keys or the values from the original dictionary: | | data | keys | values | | - | ---------------- | -------- | -------- | | 0 | {'a': 1, 'b': 2} | \[a, b\] | \[1, 2\] | | 1 | {'a': 3, 'b': 4} | \[a, b\] | \[3, 4\] | ## Summary In this article, we looked at different solutions for extraction and explosion of dictionary columns in Pandas. We focused on extraction in these code snippets, but exploding and splitting is very similar. ### How to Convert DataFrame to JSON without Backslash in Pandas URL: https://datascientyst.com/convert-dataframe-to-json-without-backslash-in-pandas/ Last updated: 2024-04-09T10:02:28.000Z In this short tutorial, you'll see the steps to convert DataFrame to JSON without backslash escape in Pandas and Python. **Note:** Read also: [How to Export DataFrame to JSON with Pandas ](https://datascientyst.com/export-dataframe-to-json-pandas/) ![](https://datascientyst.com/content/images/2023/03/convert-dataframe-to-json-without-backslash-in-pandas.webp) Suppose we have the following DataFrame: | | Name | Age | site | | - | ------- | --- | ------------------- | | 0 | Alice | 25 | http://example.com/ | | 1 | Bob | 30 | http://example.com/ | | 2 | Charlie | 35 | http://example.com/ | Here is the result of the conversion with `to_json()` with option `orient='records'`: ```python df.to_json(orient='records',lines=True) ``` We get a valid JSON file but with extract backslashes: ``` {"Name":"Alice","Age":25,"site":"http:\/\/example.com\/"} {"Name":"Bob","Age":30,"site":"http:\/\/example.com\/"} {"Name":"Charlie","Age":35,"site":"http:\/\/example.com\/"} ``` To convert Pandas DataFrame to JSON file without backslash escapes: ```python formatted_json = df.to_json(orient='records',lines=True).replace('\\/', '/') print(formatted_json) ``` This will replace all additional backslash escapes with: ``` {"Name":"Alice","Age":25,"site":"http://example.com/"} {"Name":"Bob","Age":30,"site":"http://example.com/"} {"Name":"Charlie","Age":35,"site":"http://example.com/"} ``` To store the JSON data as a file without backslash we can do: ```python print(formatted_json, file=open('data.json', 'w')) ``` We can get this result without `lines=True`: ``` [ { "Name": "Alice", "Age": 25, "site": "http://example.com/" }, { "Name": "Bob", "Age": 30, "site": "http://example.com/" }, { "Name": "Charlie", "Age": 35, "site": "http://example.com/" } ] ``` ### How to Convert Pandas Column or Row to List URL: https://datascientyst.com/how-to-convert-pandas-column-or-row-to-list/ Last updated: 2023-03-17T21:21:25.000Z To convert a DataFrame column or row to a list in Pandas, we can use the Series method `tolist()`. Here's how to do it: ```python df['A'].tolist() df.B.tolist() ``` Image below shows some of the solutions described in this article: ![](https://datascientyst.com/content/images/2023/03/convert-pandas-column-or-row-to-list.webp) ## Setup We will use the following DataFrame to convert rows and columns to list: ```python ```python import pandas as pd df = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]}) ``` data: | | A | B | | - | - | - | | 0 | 1 | 4 | | 1 | 2 | 5 | | 2 | 3 | 6 | ## Convert column to list To convert a column to a list, we can access the column using either - the bracket notation `([])` - or the dot notation `(.)` and then call the `tolist()` method on the resulting pandas Series object: ```python df['A'].tolist() df.B.tolist() ``` result: ``` [1, 2, 3] [4, 5, 6] ``` ## Convert row to list To get a list from a row in a Pandas DataFrame, we can use the `iloc` indexer to access the row by its index. Then call the `tolist()` method on the resulting pandas Series object: ```python row_list = df.iloc[0].tolist() ``` result: ``` [1, 4] ``` ## Convert column with list values to row We can convert columns which have list values to rows by using method `.explode()`. The example below shows how to convert column B which has list values. We will use new DataFrame: ```python df = pd.DataFrame({'A':['a','b'], 'B':[['1', '2'],['3', '4', '5']]}) ``` with data: | | A | B | | - | - | ----------- | | 0 | a | \[1, 2\] | | 1 | b | \[3, 4, 5\] | To rows by: ```python df.explode('B') ``` | | A | B | | - | - | - | | 0 | a | 1 | | 0 | a | 2 | | 1 | b | 3 | | 1 | b | 4 | | 1 | b | 5 | ## Column from lists (string) to lists Finally let's check how to convert column which contains list values stored as strings to list: ```python import pandas as pd data_dict = {'one': pd.Series([1, 2, 3], index=['a', 'b', 'c']), 'two': pd.Series(['[1, 2]', '[3, 4]', '[5, 6]'], index=['a', 'b', 'c'])} df = pd.DataFrame(data_dict) ``` data looks like: | | one | two | | - | --- | -------- | | a | 1 | \[1, 2\] | | b | 2 | \[3, 4\] | | c | 3 | \[5, 6\] | We can extract list values by using list comprehensions: ```python [x.strip('[]').split(',') for x in df['two']] ``` which result into new list of lists: ``` [['1', ' 2'], ['3', ' 4'], ['5', ' 6']] ``` If we like to keep the list values in the column we can use `ast` module: ```python import ast df.two.apply(ast.literal_eval) ``` Which we result into: ``` a [1, 2] b [3, 4] c [5, 6] Name: two, dtype: object ``` ### Pandas vs Julia - cheat sheet and comparison URL: https://datascientyst.com/pandas-vs-julia-comparison-cheat-sheet/ Last updated: 2023-12-01T13:39:19.000Z This is a **Python/Pandas vs Julia cheatsheet and comparison**. You can find what is the **equivalent of Pandas in Julia** or vice versa. You can find links to the documentation and other useful Pandas/Julia resources. The table below show the useful links for both: | | Pandas | Julia | | -------- | -------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- | | | data analysis tool | high performance language | | site | [https://pandas.pydata.org/](https://pandas.pydata.org/?ref=datascientyst.com) | [https://julialang.org/](https://julialang.org/?ref=datascientyst.com) | | docs | [https://pandas.pydata.org/docs/](https://pandas.pydata.org/docs/?ref=datascientyst.com) | [https://docs.julialang.org/en/v1/](https://docs.julialang.org/en/v1/?ref=datascientyst.com) | | packages | [https://pypi.org/](https://pypi.org/?ref=datascientyst.com) | [https://juliapackages.com/](https://juliapackages.com/?ref=datascientyst.com) | | repo | [https://github.com/pandas-dev/pandas](https://github.com/pandas-dev/pandas?ref=datascientyst.com) | [https://github.com/JuliaLang/julia](https://github.com/JuliaLang/julia?ref=datascientyst.com) | Below you can find equivalent code between Pandas and Julia. Have in mind that some examples might differ due to different indexing. Single column 2 columns 3 columns Hide navigation Hide TOC ## Setup Import and package installation ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_basics.png) `import pandas as pd import numpy as np` using DataFrames using Statistics using CSV Import libraries and modules `pip install pandas` using Pkg Pkg.add("JSON") install package `https://pypi.org/` https://juliapackages.com/ Search Packages ## Data Structures Pandas Series vs Julia Array DataFrame comparison ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_data_structures-1.png) `s = pd.Series(['a', 'b', 'c'], index=[0 , 1, 2])` s = \[1, 2, 3\] Pandas series vs Julia vector `s[0]` s\[1\] Get first element of array or Series `df = pd.DataFrame( {'col_1': [11, 12, 13], 'col_2': [21, 22, 23]}, index=[0, 1, 3])` df = DataFrame(a=11:13, b=21:23) Pandas vs Julia DataFrame `import numpy as np import pandas as pd data=np.random.randint(0,10,size=(10, 3)) df = pd.DataFrame(data, columns=list('abc'))` using Random Random.seed!(1); df = DataFrame(rand(10, 3), \[:a, :b, :c\]) Create random DataFrame ## Read Import Data Julia vs Pandas ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_read.png) `df = pd.read_csv('file.csv')` df = CSV.read("file.csv", DataFrame) Read CSV file `pd.read_json('file.json')` using JSON JSON.parsefile("file.json") Read JSON file `pd.read_csv('https://example.com/file.csv')` A = urldownload("https://example.com/file.csv") A |> DataFrame Read data from URL `df = pd.read_fwf('delim_file.txt')` readdlm("delim\_file.txt", ' ', Int, ' ') Read delimited file ## Write Data export - Pandas vs Julia ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_write.png) `df.to_csv('file.csv')` CSV.write("file.csv", df) Writes to a CSV file `df.to_json(filename)` using JSON3 JSON3.write("file.json",df1) Writes to a file in JSON format ## Inspect Data Statistics, samples and summary of the data ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_info.png) `df.head(6)` first(df, 6) First n rows `df.tail(6)` last(df, 6) Last n rows `df.describe()` describe(df) Summary statistics `df.loc[:, :'a'].describe()` describe(df\[!, \[:a\]\]) Describe columns `df['A'].mean()` using Statistics mean(df.A) Statistical functions ## Select Select data by index, by label, get subset ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_select.png) `df.loc[1:3, :]` df\[1:3, :\] Select first N rows - all columns `df.loc[[1, 2, 3], :]` df\[\[1, 2, 3\], :\] Select rows by index `df.loc[:, ['a', 'b']].copy()` df\[:, \[:a, :b\]\] Select columns by name(copy) `df.loc[:, ['a']]` df\[!, \[:A\]\] Select columns by name(reference) `df.loc[1:3, ['b', 'a']]` df\[1:3, \[:b, :a\]\] Subset rows and columns `df.loc[[3,1], ['b', 'a']]` df\[\[3, 1\], \[:c\]\] Reverse selection `df[df['a'].isna()]` findall(ismissing, df\[:, "a"\]) Select NaN values `df['a'].dropna()` filter(!ismissing, df\[:, "a"\]) Select non NaN values ## Add rows/columns Add new columns and rows ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_add.png) `df['new col'] = df['col'] * 100` df\[!, "d"\] = df\[!, "a"\] \* 100 Add new column based on other column `df['new col'] = False` df\[!, "e"\] .= false Add new column single value `df.loc[-1] = [1, 2, 3]` push!(df,\[0, 0, 0\]) Add new row at the end of DataFrame `df.append(df2, ignore_index = True)` append!(df,df2) add rows from DataFrame to existing DataFrame ## Drop rows/columns/nan Drop data from DataFrame ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_drop.png) `s.drop(1)` filter!(e->e≠1,a) (Series) Drop values from Series by index (row axis) `s.drop([1, 2])` filter!(e->e∉\[1, 2\],a) (Series) Drop values from Series by index (row axis) `df.drop('b' , axis=1) ` dropmissing!(df\[:, \["b"\]\]) Drop column by name col\_1 (column axis) `df.dropna()` dropmissing!(df) Drops all rows that contain null values `df.dropna()` df\[all.(!ismissing, eachrow(df)), :\] Drops all rows that contain null values `df.dropna(axis=1)` df\[:, all.(!ismissing, eachcol(df))\] Drops all columns that contain null values ## Sort values/index Sorting and rank values in Pandas vs Julia ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_sort.png) `sorted([2,3,1])` sort(\[2,3,1\]) sort array of values `sorted([2,3,1], reverse=True)` sort(\[2,3,1\], rev=true) sort in reverse order `df['a'].sort_values()` sort(df, \[:a\]) sort DataFrame by column `df.sort_values(['a', 'b'], ascending=[False, True])` sort(df, \[order(:a, rev=true), :b\]) sort DataFrame by multiple columns ## Filter Filter data based on multiple criteria ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_filter.png) `df.loc[:, df.isna().any()]` mapcols(x -> any(ismissing, x), df) find columns with na `df[df['col_1'] > 100]` filter(row -> row.a > 100, df) Values greater than X `df[(df['a']=='a')&(df['b']>=10)]` filter(row -> row.a == 'a' && row.b >= 5, df) Filter Multiple Conditions - & - and; | - or `df[df['a'] == 'test']` df\[ ( df.a .== "test" ) , :\] filter by sting value `df[(df['a'] == 'test') & (df['b'] == 'a2') ]` df\[ ( df.a .== "test" ) .& ( df.b .== "a2" ), :\] combine conditions ## Group by Group by and summarize data ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_groupby.png) `df.groupby('a')` groupby(df, \[:a\]) Group by single column `df.groupby(['a', 'b']).c.sum()` gdf = groupby(df, \[:a, :b\]) combine(gdf, :c => sum) group by multiple columns and sum third `df['a'].value_counts()` combine(groupby(df, \[:x1\]), nrow => :count) group by and count ## Convert Convert to date, string, numeric ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_convert.png) `df['a'].fillna(0)` replace(df.a,missing => 0) replace NA values `df.replace('..', None)` ifelse.(df .== "..", missing, df) convert .. to NA `df['col_1'].astype('int64')` df\[!, :a\] = parse.(Int64, df\[!, :a\]) convert string to int `pd.to_datetime(df['date'], format='%Y-%m-%d')` using Dates df.Date = Date.(df.Date, "dd-mm-yyyy") convert string to date ## Install Julia Packages To install new packages in Julia we can also use the Julia Package manager by: - open Linux Terminal - start Julia - `julia` - Type `]` (right bracket). You don’t have to hit Return. - Termimal will change to `(@v1.8) pkg>` - Type add to add a package - you can provide the names of several packages separated by spaces. - `Control-C` to exit the package manager Example: `(v1.8) pkg> add JSON StaticArrays` ## Differences: Julia and Pandas Pandas and Julia are both popular tools for data analysis and manipulation. Some key differences between them: ### Indexing One big difference between Julia and Pandas is indexing: - Julia - 1 based (can be configured) \* [What's the big deal? 0 vs 1 based indexing](https://discourse.julialang.org/t/whats-the-big-deal-0-vs-1-based-indexing/1102/13?ref=datascientyst.com) \* [Why does Julia adopt 1-based index?](https://www.reddit.com/r/Julia/comments/pm1yl9/why%5Fdoes%5Fjulia%5Fadopt%5F1based%5Findex/?ref=datascientyst.com) \* Pandas - 0 based ### Syntax Personally I prefer SQL syntax over both Julia and Pandas. I can work fine with both of them. As I have more experience with Python I would go with Pandas. Some people consider Julia to have better syntax since it was designed for data science. Example of syntax difference between Julia and Pandas: ```python # pandas import pandas as pd df = pd.read_csv('sales_data.csv') totals = df.groupby('product')['sales'].sum() # julia using DataFrames using CSV df = DataFrame(CSV.read("sales_data.csv")) totals = combine(groupby(df, :product), :sales => sum) ``` ### Performance In general Julia is faster for most operations and bigger datasets. For smaller datasets Pandas might be close or even better than Julia. The reason is for compilation time for Julia. To test performance we can use dataset with 10M rows - [Game Recommendations on Steam](https://www.kaggle.com/datasets/antonkozyriev/game-recommendations-on-steam?select=recommendations.csv&ref=datascientyst.com): ```python # pandas %%time import pandas as pd df = pd.read_csv('recommendations.csv') df['hours'].mean() # julia @time begin using CSV, DataFrames df = CSV.File("recommendations.csv") |> DataFrame result = mean(df[:, "hours"]) end ``` The results are: - Pandas - CPU times: user 5.67 s, sys: 1.85 s, total: 7.52 s - Wall time: 7.71 s - Julia - 7.257497 seconds (1.13 k allocations: 1.349 GiB, 2.63% gc time) While for [dataset](https://www.kaggle.com/datasets/lasaljaywardena/12-million-company-data?ref=datascientyst.com) \- 12M rows we get: - Pandas - CPU times: user 34.8 s, sys: 3.74 s, total: 38.5 s - Wall time: 42.4 s - Julia - 29.964544 seconds (162.12 M allocations: 9.878 GiB, 15.79% gc time) First julia execution is slower so we take the second one. ### Libraries and Ecosystem Pandas has a bigger community and ecosystem. The Python libraries offers greater variety of Packages in many areas: - web scraping - data science - science - etc ### Language Features I prefer Julia for distributed computing and parallel computing. Pandas seems for me much better for visualization and EDA. ### Learning Curve Again it depends on personal choice. Python is considered as one of the best programming languages for beginners. Julia surpassed Python in recent surveys for loved language: [stackoverflow survey - Most loved, dreaded, and wanted](https://survey.stackoverflow.co/2022/?ref=datascientyst.com#technology-most-loved-dreaded-and-wanted) Note: I need to add that I'm still learning and discovering Julia - so so statements above might change in future :) ## Pandas vs Julia docs | | Pandas | Julia | | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | | NA | NA | missing | | Boolean | False/True | false/true | | docs | [https://pandas.pydata.org/docs/](https://pandas.pydata.org/docs/?ref=datascientyst.com) | [https://docs.julialang.org/en/v1/](https://docs.julialang.org/en/v1/?ref=datascientyst.com) | | packages | [https://pypi.org/](https://pypi.org/?ref=datascientyst.com) | [https://juliapackages.com/](https://juliapackages.com/?ref=datascientyst.com) | | repo | [https://github.com/pandas-dev/pandas](https://github.com/pandas-dev/pandas?ref=datascientyst.com) | [https://github.com/JuliaLang/julia](https://github.com/JuliaLang/julia?ref=datascientyst.com) | | basics | [https://pandas.pydata.org/docs/user\_guide/basics.html](https://pandas.pydata.org/docs/user%5Fguide/basics.html?ref=datascientyst.com) | [https://docs.julialang.org/en/v1/base/punctuation/](https://docs.julialang.org/en/v1/base/punctuation/?ref=datascientyst.com) | | start | [https://pandas.pydata.org/docs/getting\_started/index.html](https://pandas.pydata.org/docs/getting%5Fstarted/index.html?ref=datascientyst.com) | [https://juliadatascience.io/](https://juliadatascience.io/?ref=datascientyst.com) | ## Summary In summary, Pandas and Julia are both powerful tools for data analysis, but they have different strengths and weaknesses. Pandas has a larger ecosystem of tools and is generally easier to learn. Julia is faster and has some unique language features that can make it more powerful for certain types of data analysis tasks. Ultimately, the choice between Pandas and Julia depends on your specific requirements and preferences. ## Resources - [Julia Comparison with the Python package Pandas](https://dataframes.juliadata.org/stable/man/comparisons/?ref=datascientyst.com#Comparison-with-the-Python-package-pandas) - [Pandas Cheat Sheet for Data Science](https://datascientyst.com/pandas-cheat-sheet-for-data-science/) - [Pandas vs SQL Cheat Sheet](https://datascientyst.com/pandas-vs-sql-cheat-sheet/) - [julia notebook](https://github.com/softhints/Pandas-Exercises-Projects/blob/main/cheat%5Fsheet/pandas%5Fjulia/julia.ipynb?ref=datascientyst.com) - [pandas notebook](https://github.com/softhints/Pandas-Exercises-Projects/blob/main/cheat%5Fsheet/pandas%5Fjulia/pandas.ipynb?ref=datascientyst.com) ## Cheatsheet Image ![](https://datascientyst.com/content/images/2023/03/pandas-vs-julia-light.webp) ### Pandas random sampling: stratified and weighted URL: https://datascientyst.com/pandas-random-sampling-stratified-and-weighted/ Last updated: 2023-02-19T10:27:55.000Z In this quick tutorial, we're going to discuss stratified sampling in Pandas and Python. The following syntax can be used to sample stratified in Pandas: **(1) stratified sampling - disproportionated** ```python (df .groupby('continent', group_keys=False) .apply(lambda x: x.sample(2)) ) ``` **(2) stratified sampling - proportional** ```python (df .groupby('continent', group_keys=False) .apply(lambda x: x.sample(frac=0.1)) ) ``` The image below illustrates the technique called stratified sampling: ![](https://datascientyst.com/content/images/2023/02/pandas-random-sampling-stratified-and-weights.webp) Next, you'll see the steps to do stratified sampling in practice. ## Setup First, let's create a sample DataFrame: ```python import plotly.express as px df = px.data.gapminder().query("year == 2007") cols = df.columns[:4] df = df[cols] ``` data has the following shape: ``` (142, 5) ``` First rows of this DataFrame | | country | continent | year | lifeExp | freq | | -- | ----------- | --------- | ---- | ------- | ---- | | 11 | Afghanistan | Asia | 2007 | 43.828 | 33 | | 23 | Albania | Europe | 2007 | 76.423 | 30 | | 35 | Algeria | Africa | 2007 | 72.301 | 52 | | 47 | Angola | Africa | 2007 | 42.731 | 52 | | 59 | Argentina | Americas | 2007 | 75.320 | 25 | ## Separate population into strata The whole population of this dataset is 142 countries. Below you can find the proportions per each stratum: ```python df[col_name].value_counts() ``` result is: ``` Africa 52 Asia 33 Europe 30 Americas 25 Oceania 2 Name: continent, dtype: int64 ``` ## Find the sample size Next we need to decide on what should be the size of the sample. There are different strategies on this. These are the options available in Pandas `sample()` method: - `n` \- number of items to return. We get N random items - `frac` \- fraction of items to return - `weights` \- probability weighting So we can select the size of: - total sample as percentage on the whole population - each group - disproportionated - proportionated ## Disproportional stratified sampling In this approach the size of each sample group is not proportional to the entire population. We will get equal number of items for each group: ```python (df .groupby('continent', group_keys=False) .apply(lambda x: x.sample(2)) ) ``` The result is 2 sized stratum - no matter the size of each group: | | country | continent | year | lifeExp | freq | | ---- | ------------------ | --------- | ---- | ------- | ---- | | 911 | Libya | Africa | 2007 | 73.952 | 52 | | 899 | Liberia | Africa | 2007 | 45.678 | 52 | | 443 | Dominican Republic | Americas | 2007 | 72.235 | 25 | | 791 | Jamaica | Americas | 2007 | 72.567 | 25 | | 1679 | Yemen, Rep. | Asia | 2007 | 62.698 | 33 | | 1319 | Saudi Arabia | Asia | 2007 | 72.777 | 33 | | 779 | Italy | Europe | 2007 | 80.546 | 30 | | 1607 | United Kingdom | Europe | 2007 | 79.425 | 30 | | 1103 | New Zealand | Oceania | 2007 | 80.204 | 2 | | 71 | Australia | Oceania | 2007 | 81.235 | 2 | ## Proportional stratified sampling Taking random sampling from stratified groups which is proportional to the population. We can do proportional stratified sampling in Pandas by sampling with parameter `x.sample(frac=0.1)`: ```python (df .groupby('continent', group_keys=False) .apply(lambda x: x.sample(frac=0.1)) ) ``` This will give us countries proportioned to the initial population: | | country | continent | year | lifeExp | freq | | ---- | ------------------ | --------- | ---- | ------- | ---- | | 491 | Equatorial Guinea | Africa | 2007 | 51.579 | 52 | | 335 | Congo, Dem. Rep. | Africa | 2007 | 46.462 | 52 | | 1691 | Zambia | Africa | 2007 | 42.384 | 52 | | 131 | Benin | Africa | 2007 | 56.728 | 52 | | 1571 | Tunisia | Africa | 2007 | 73.923 | 52 | | 443 | Dominican Republic | Americas | 2007 | 72.235 | 25 | | 1643 | Venezuela | Americas | 2007 | 73.747 | 25 | | 815 | Jordan | Asia | 2007 | 72.535 | 33 | | 1655 | Vietnam | Asia | 2007 | 74.249 | 33 | | 875 | Lebanon | Asia | 2007 | 71.993 | 33 | | 1607 | United Kingdom | Europe | 2007 | 79.425 | 30 | | 1091 | Netherlands | Europe | 2007 | 79.762 | 30 | | 539 | France | Europe | 2007 | 80.657 | 30 | Africa is the most represented continent while Oceania is missing from this sample. ### color row based on column value If you like to learn how to style each group into different color check: [color Pandas DataFrame based on value](https://datascientyst.com/pandas-dataframe-background-color-based-condition-value-alternate-row-color-based-group/) ```python def format_color_groups(df): colors = ['gold', 'lightblue'] x = df.copy() factors = list(x['continent'].unique()) i = 0 for factor in factors: style = f'background-color: {colors[i]}' x.loc[x['continent'] == factor, :] = style i = not i return x d1.style.apply(format_color_groups, axis=None) ``` The result is: ![pandas-random-sampling-stratified-proprotional](https://datascientyst.com/content/images/2023/02/pandas-random-sampling-stratified-proprotional.webp) ## group by sample weights in Pandas We can use weights to get proportional sampling. First we need to calculate the weights of each stratum: ### Calc weights - groupby + .transform('count') ```python df['weight'] = (df .groupby('continent') .country .transform('count') ) ``` as a result we have Pandas series which contains the weight for each row: ``` 11 33 23 30 35 52 47 52 59 25 .. 1655 33 1667 33 1679 33 1691 52 1703 52 Name: country, Length: 142, dtype: int64 ``` ### Calc weights - equal representation To get equal representation of each group we can calculate the weights: ```python df['weight'] = 1./(df .groupby('continent') .country .transform('count') ) ``` The get disproportional rate: ``` 11 0.030303 23 0.033333 35 0.019231 47 0.019231 59 0.040000 ... 1655 0.030303 1667 0.030303 1679 0.030303 1691 0.019231 1703 0.019231 Name: weight, Length: 142, dtype: float64 ``` ### Calc weights - map and value\_counts We can achieve the same result by `.map` and `.value_counts()` ```python df['weight'] = (df .continent .map(df['continent'].value_counts()) ) ``` ### sampling with weights To do sampling with respect to distribution of a value in a given column we can use the calculate frequency for parameter `weights`: ```python df.sample(n=10, weights = df['weight']) ``` Result is weighted sampling: | | country | continent | year | lifeExp | freq | weight | | ---- | ------------------- | --------- | ---- | ------- | ---- | ------ | | 851 | Korea, Rep. | Asia | 2007 | 78.623 | 33 | 33 | | 491 | Equatorial Guinea | Africa | 2007 | 51.579 | 52 | 52 | | 1043 | Mozambique | Africa | 2007 | 42.082 | 52 | 52 | | 1175 | Pakistan | Asia | 2007 | 65.483 | 33 | 33 | | 1223 | Philippines | Asia | 2007 | 71.688 | 33 | 33 | | 1667 | West Bank and Gaza | Asia | 2007 | 73.422 | 33 | 33 | | 1031 | Morocco | Africa | 2007 | 71.164 | 52 | 52 | | 1139 | Nigeria | Africa | 2007 | 46.859 | 52 | 52 | | 1559 | Trinidad and Tobago | Americas | 2007 | 69.819 | 25 | 25 | | 635 | Guinea-Bissau | Africa | 2007 | 46.388 | 52 | 52 | ## Conclusion In this article, we took a closer look at stratified sampling in Pandas and how to apply it in practice. We covered the two main approaches in stratified sampling - disproportionated and proportionated. Finally we explain how to use weights to sample and groupby in Pandas DataFrame. ## resources - [Stratified sampling](https://en.wikipedia.org/wiki/Stratified%5Fsampling?ref=datascientyst.com) - [Random Sample per group in pandas](https://datascientyst.com/random-sample-per-group-in-pandas/) ### Random Sample per group in pandas URL: https://datascientyst.com/random-sample-per-group-in-pandas/ Last updated: 2023-03-15T15:51:46.000Z Here are several ways to **sample random rows per group in Pandas**: **(1) random selection per group** ```python df.groupby('continent').apply(lambda x: x.sample(n=3)) ``` **(2) random selection per group - different size** ```python (df .groupby('continent') .apply(lambda x: x.sample(n=3, replace=True)) .drop_duplicates() ) ``` **(3) sample based on column** ```python col = 'continent' categories = list(df[col].dropna().unique()) for cat in categories: if df[df[col] == cat].shape[0] > 3: print(cat, end=' - ') display(df[df[col] == cat].sample(3, replace=True)) ``` The result is shown below - getting N random samples per each group on: ![](https://datascientyst.com/content/images/2023/02/random-sample-per-group-in-pandas.webp) ## Setup In the post, we'll use the following DataFrame, available from library `plotly`. To install plotly use `pip install plotly`. We will create DataFrame getting info for 2007 year: ```python import plotly.express as px df = px.data.gapminder().query("year == 2007") cols = df.columns[:4] df = df[cols] ``` DataFrame looks like: | | country | continent | year | lifeExp | | -- | ----------- | --------- | ---- | ------- | | 11 | Afghanistan | Asia | 2007 | 43.828 | | 23 | Albania | Europe | 2007 | 76.423 | | 35 | Algeria | Africa | 2007 | 72.301 | | 47 | Angola | Africa | 2007 | 42.731 | | 59 | Argentina | Americas | 2007 | 75.320 | For simplicity we will work only with the first 4 columns. ## 1: Random selection per group To do random selection per group in Pandas we can: - use `groupby()` on a column(s) - and use `apply` and `sample` methods: ```python df.groupby('continent').apply(lambda x: x.sample(n=1)) ``` We get random country per each continent: | | | country | continent | year | lifeExp | | --------- | ---- | --------- | --------- | ---- | ------- | | continent | | | | | | | Africa | 1595 | Uganda | Africa | 2007 | 51.542 | | Americas | 1211 | Peru | Americas | 2007 | 71.421 | | Asia | 719 | Indonesia | Asia | 2007 | 70.650 | | Europe | 527 | Finland | Europe | 2007 | 79.313 | | Oceania | 71 | Australia | Oceania | 2007 | 81.235 | If we try to use large sample number we will get error: > ValueError: Cannot take a larger sample than population when 'replace=False' We will solve this error in next section ## sample by group - different size If the groups have different sizes or some groups are smaller than the selected sample number we can use repetitions. To sample with repetition we need to pass `replace=True` to `sample()` method: ```python df.groupby('continent').apply(lambda x: x.sample(n=3, replace=True)) ``` This will solve the error above: | | | country | continent | year | lifeExp | | --------- | -------------- | ----------- | --------- | ------ | ------- | | continent | | | | | | | Africa | 1547 | Togo | Africa | 2007 | 58.420 | | 347 | Congo, Rep. | Africa | 2007 | 55.322 | | | 1571 | Tunisia | Africa | 2007 | 73.923 | | | Americas | 251 | Canada | Americas | 2007 | 80.653 | | 287 | Chile | Americas | 2007 | 78.553 | | | 251 | Canada | Americas | 2007 | 80.653 | | | Asia | 1439 | Sri Lanka | Asia | 2007 | 72.396 | | 731 | Iran | Asia | 2007 | 70.964 | | | 107 | Bangladesh | Asia | 2007 | 64.062 | | | Europe | 527 | Finland | Europe | 2007 | 79.313 | | 407 | Czech Republic | Europe | 2007 | 76.486 | | | 1283 | Romania | Europe | 2007 | 72.476 | | | Oceania | 1103 | New Zealand | Oceania | 2007 | 80.204 | | 71 | Australia | Oceania | 2007 | 81.235 | | | 1103 | New Zealand | Oceania | 2007 | 80.204 | | ### Get unique samples per group Downside is that we have repetitions in the groups with smaller items. Remove duplicated rows can be done by adding `.drop_duplicates()`: ```python (df .groupby('continent') .apply(lambda x: x.sample(n=3, replace=True)) .drop_duplicates() ) ``` This code shows the usage of Pandas chaining. It's easier to read and maintain. ## Group by multiple columns and sample To group by multiple columns and sample random values per each group in Pandas we can use similar code: ```python (df .groupby(['continent', 'year']) .apply(lambda x: x.sample(n=2, replace=True)) .drop_duplicates() ) ``` The result is multi-index with samples from each group: | | | | country | continent | year | lifeExp | | --------- | ----------- | -------- | --------- | --------- | ---- | ------- | | continent | year | | | | | | | Africa | 2007 | 1067 | Namibia | Africa | 2007 | 52.906 | | 923 | Madagascar | Africa | 2007 | 59.443 | | | | Americas | 2007 | 179 | Brazil | Americas | 2007 | 72.390 | | 1259 | Puerto Rico | Americas | 2007 | 78.746 | | | | Asia | 2007 | 1007 | Mongolia | Asia | 2007 | 66.803 | | 875 | Lebanon | Asia | 2007 | 71.993 | | | | Europe | 2007 | 1391 | Slovenia | Europe | 2007 | 77.926 | | 1091 | Netherlands | Europe | 2007 | 79.762 | | | | Oceania | 2007 | 71 | Australia | Oceania | 2007 | 81.235 | | 1103 | New Zealand | Oceania | 2007 | 80.204 | | | ## Random sample per group - for loop If we need to add filtering or parsing logic we may use a `for` loop. In this way we may get different DataFrames for each sample group: ```python col = 'continent' categories = list(df[col].dropna().unique()) for cat in categories: group_size = df[df[col] == cat].shape[0] print(cat, '-', group_size) if group_size >= 3: display(df[df[col] == cat].sample(3)) else: display(df[df[col] == cat].sample(group_size)) ``` The result is visible on the image below: ![random-sample-per-group-in-pandas-python](https://datascientyst.com/content/images/2023/02/random-sample-per-group-in-pandas-python.webp) ## Conclusion In this article, we looked at different ways for sampling random items per group in Pandas and Python. We focused on equal sampling, but also covered sampling from groups with different sizes. Finally we saw how to exclude, filter or parse some groups when doing random sampling. We briefly introduce Pandas chaining - a nice technique for writing better and more readable Pandas code. ## Resources - [Pandas - Random Sample of a subset of a DataFrame](https://softhints.com/pandas-random-sample-of-a-subset-of-a-dataframe-rows-or-columns?ref=datascientyst.com) - [How to Get Top 10 Highest or Lowest Values in Pandas](https://datascientyst.com/get-top-10-highest-lowest-values-pandas/) - [Pandas random sampling: stratified and weighted](https://datascientyst.com/pandas-random-sampling-stratified-and-weighted/) ### How to Convert List of Objects to Pandas DataFrame? URL: https://datascientyst.com/convert-list-of-objects-to-pandas-dataframe/ Last updated: 2023-02-18T08:59:31.000Z To **convert a list of objects to a Pandas DataFrame**, we can use the: - `pd.DataFrame` constructor - method `from_records()` and list comprehension: **(1) Define custom class method** ```python pd.DataFrame([p.to_dict() for p in persons]) ``` **(2) Use vars() function** ```python pd.DataFrame([vars(p) for p in persons]) ``` **(3) Use attribute dict** ```python pd.DataFrame([p.__dict__ for p in persons]) ``` Here are the general steps you can follow: - inspect the objects and the class definition - convert list of objects based on the class Let's check the steps to convert a list of objects in more detail. You can find visual summary of the article in this image: ![python list of objects to pandas dataframe](https://datascientyst.com/content/images/2023/02/convert-list-of-objects-to-pandas-dataframe.webp) ## Setup To start, create your Python class - Person: ```python class Person: def __init__(self, name, age, gender): self.name = name self.age = age self.gender = gender def to_dict(self): return { "name": self.name, "age": self.age, "gender": self.gender } ``` Let’s create the following 3 objects and list of them: ```python p1 = Person("Alice", 25, "Female") p2 = Person("John", 25, "Male") p3 = Person("Tim", 30, "Male") persons = [p1, p2, p3] ``` ## 1: Convert list of objects - user method We can convert a list of model objects to Pandas DataFrame by defining a custom method. This will avoid potential errors and it's a good practice to follow. In this way we have full control on the conversion. We will use Python list comprehension for the conversion: ```python pd.DataFrame([p.to_dict() for p in persons]) ``` result is: | | name | age | gender | | - | ----- | --- | ------ | | 0 | Alice | 25 | Female | | 1 | John | 25 | Male | | 2 | Tim | 30 | Male | Conversion mapping can be changed from the method `to_dict()`: ```python def to_dict(self): return { "name": self.name, "age": self.age, "gender": self.gender } ``` ## 2: attr **dict** \- list of objects to dataframe Sometimes we don't have control of the class. In this case we may use the Python attribute `__dict__` to convert the objects to dictionaries. Once we have a list of dictionaries we can create DataFrame. So we use list comprehension and convert each object to dictionary: ```python pd.DataFrame([p.__dict__ for p in persons]) ``` the result is the same as before: | | name | age | gender | | - | ----- | --- | ------ | | 0 | Alice | 25 | Female | | 1 | John | 25 | Male | | 2 | Tim | 30 | Male | Disadvantages of this way are potential errors due to incorrect mapping or complex data types. ## 3: vars() - convert object list to dataframe The Python `vars()` function returns the **dict** attribute of an object. So this way is pretty similar to the previous. This way is more pythonic and easier to read. ```python pd.DataFrame([vars(p) for p in persons]) ``` We got the same result. Which one to choose is personal choice. I prefer `vars()` because I use: `len(my_list)` and not `my_list.__len__()`. ## 4: `from_records()` vs pd.DataFrame To convert list of objects or dictionaries can also use method `from_records()`: ```python pd.DataFrame.from_records([p.to_dict() for p in persons]) ``` In the example above the usage of both will be equivalent. The difference is the parameters for both: The parameters for `pd.DataFrame` are limited to: - data - index - columns - dtype - copy While by using `from_records()` we have better control on the conversion by option `orient`: - ‘columns’ - ‘index’ - ‘tight’ where > The “orientation” of the data. If the keys of the passed dict should be the columns of the resulting DataFrame, pass ‘columns’ (default). Otherwise if the keys should be rows, pass ‘index’. If ‘tight’, assume a dict with keys \[‘index’, ‘columns’, ‘data’, ‘index\_names’, ‘column\_names’\]. We can also create a multiindex dataFrame from a dictionary - you can read more on: [How to Create DataFrame from Dictionary in Pandas?](https://datascientyst.com/create-dataframe-from-dictionary-pandas/) ## 5\. parse objects in for loop We can use for loop to define conversion logic: ```python rows = [] for p in persons: row = { "name": p.name, "age": p.age, "gender": p.gender } rows.append(row) df = pd.DataFrame(rows) ``` In this way we iterate over objects and extract only fields of interest for us. ## 6\. JSON serializable objects To convert JSON serializable objects to Pandas DataFrame we can use: ```python import json json.dumps(p1) ``` or: ```python json.dumps(person1, default=vars) ``` which will give us: ``` '{"name": "Alice", "age": 25, "gender": "Female"}' ``` ## Conclusion In this post, we saw how to **convert a list of Python objects to Pandas DataFrame**. We covered conversion with class methods and using built-in functions. Examples with object parsing and JSON serializable objects were shown. Finally we discussed which way is better - Pandas DataFrame constructor or by `from_dict()`. ## Resource - [pandas.DataFrame.from\_dict](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.from%5Fdict.html?ref=datascientyst.com) - [pandas.DataFrame](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.html?ref=datascientyst.com) - [How to Create DataFrame from Dictionary in Pandas?](https://datascientyst.com/create-dataframe-from-dictionary-pandas/) ### How to apply Formatting and Borders to Pivot Table in Pandas URL: https://datascientyst.com/how-to-apply-formatting-and-borders-to-pivot-table-in-pandas/ Last updated: 2026-01-23T17:12:26.000Z To **apply formatting and add borders to pivot tables in Pandas** we can use `style` and `set_table_styles`. You can find basic example on adding borders and formatting: - adding borders - coloring numbers based on values - format NaN and float precision ```python df_pivot.style. \ background_gradient(cmap='Reds', axis=None). \ set_table_styles( [{'selector': 'th,td,tr', 'props': [('border-style', 'solid'), ('border-width', '1px')]}]). \ format(na_rep='', precision=2) ``` The result with formatting and borders of the pivot table is below: ![add-borders-to-pivot-table-in-pandas](https://datascientyst.com/content/images/2023/02/add-borders-to-pivot-table-in-pandas.png) In this short post, we'll see several examples of applying formatting and adding borders to Pandas DataFrame. ## Setup Let's say that we have the following data: ```python import pandas as pd data = {"A": ["foo", "foo", "foo", "foo", "foo", "bar", "bar", "bar", "bar"], "B": ["one", "one", "one", "two", "two", "one", "one", "two", "two"], "C": ["small", "large", "large", "small", "small", "large", "small", "small", "large"], "D": [1, 2, 2, 3, 3, 4, 5, 6, 7], "E": [2, 4, 5, 5, 6, 6, 8, 9, 9]} df = pd.DataFrame(data) ``` DataFrame looks like: | | A | B | C | D | E | | - | --- | --- | ----- | - | - | | 0 | foo | one | small | 1 | 2 | | 1 | foo | one | large | 2 | 4 | | 2 | foo | one | large | 2 | 5 | | 3 | foo | two | small | 3 | 5 | | 4 | foo | two | small | 3 | 6 | | 5 | bar | one | large | 4 | 6 | | 6 | bar | one | small | 5 | 8 | | 7 | bar | two | small | 6 | 9 | | 8 | bar | two | large | 7 | 9 | we will create pivot table by: ```python df_pivot = pd.pivot_table(df, values=['D', 'E'], index=['A', 'C'], aggfunc={'D': ['mean', 'count', 'sum'], 'E': [min, max]}) df_pivot ``` Check this to learn more about: [How To Create a Pivot Table in Pandas](https://datascientyst.com/how-to-create-a-pivot-table-in-pandas/) ## Apply formatting to pivot table To **apply formatting to a pivoted DataFrame in Pandas**, we can use the `style` property. In addition we can use the functions: - `background_gradient` - `format` - `highlight_null` - `set_caption` to control the formatting of the DataFrame. Here's an example of how to apply: - color gradient - summer - highlight NaN values in red: - add title to pivot DataFrame - use brown for the column names ```python df_pivot.style.background_gradient(cmap='summer')\ .format('{:.2f}')\ .highlight_null(null_color='red')\ .set_table_styles([{'selector': 'th', 'props': [('color', 'brown')]}])\ .set_caption('Pivot with formatting') ``` The pivot table with formatting looks like: ![apply-formatting-to-pivot-table-in-pandas](https://datascientyst.com/content/images/2023/02/apply-formatting-to-pivot-table-in-pandas.png) ## Add borders to pivot table Let's say that we would like to **add borders to the pivot table**. We are going to use `style` and `set_table_styles`. We can use CSS selectors: - `th` \- header cell in a table - `tr` \- row in a table - `td` \- cell in a table In this example you can find different border styles applied on each selector: ```python df_pivot.style. \ set_table_styles([{'selector': 'tr', 'props': [('border', '4px solid blue')]}, {'selector': 'th', 'props': [('border', '3px solid red')]}, {'selector': 'td', 'props': [('border', '5px solid black')]}]) ``` Running the code will give as this table: ![add-borders-to-pivoted-dataframe-pandas](https://datascientyst.com/content/images/2023/02/add-borders-to-pivoted-dataframe-pandas.png) ## Summary In this post, we saw how to apply formatting, table styles and borders to pivot tables in Pandas. The correct formatting and borders will make your pivot tables easier to digest. As you can see from the image below by default pivot formatting makes the table hard to read: ![](https://datascientyst.com/content/images/2023/02/how-to-apply-formatting-and-borders-to-pivot-table-in-pandas.webp) ## Resources - [Formatting tables in Pandas](https://datascientyst.com/tag/422-table/) - [Pandas styling](https://datascientyst.com/tag/421-styling/) - [Pandas Styling tips](https://datascientyst.com/tag/420-basic-concepts/) - [How to Display Pandas DataFrame As a Heatmap](https://datascientyst.com/display-pandas-dataframe-heatmap/) ### How To Create a Pivot Table in Pandas? URL: https://datascientyst.com/how-to-create-a-pivot-table-in-pandas/ Last updated: 2023-09-07T07:01:35.000Z We can use the following syntax to **create a pivot table in Python using Pandas**: ```python df_pivot = df.pivot_table(values='D', index=['A', 'B'], columns='C') ``` Next, we'll see the full steps to create pivot tables in Pandas using a simple example. ![](https://datascientyst.com/content/images/2023/02/how-to-create-a-pivot-table-in-pandas.webp) Steps to create pivot table: ## Step 1: Get data for pivot Suppose we have the following DataFrame which contains 4 columns: - 3 string columns - 2 numeric We will pivot on multiple columns in next sections ```python import pandas as pd data = {"A": ["foo", "foo", "foo", "foo", "foo", "bar", "bar", "bar", "bar"], "B": ["one", "one", "one", "two", "two", "one", "one", "two", "two"], "C": ["small", "large", "large", "small", "small", "large", "small", "small", "large"], "D": [1, 2, 2, 3, 3, 4, 5, 6, 7], "E": [2, 4, 5, 5, 6, 6, 8, 9, 9]} df = pd.DataFrame(data) ``` Let's see how to pivot based on data below: | | A | B | C | D | E | | - | --- | --- | ----- | - | - | | 0 | foo | one | small | 1 | 2 | | 1 | foo | one | large | 2 | 4 | | 2 | foo | one | large | 2 | 5 | | 3 | foo | two | small | 3 | 5 | | 4 | foo | two | small | 3 | 6 | | 5 | bar | one | large | 4 | 6 | | 6 | bar | one | small | 5 | 8 | | 7 | bar | two | small | 6 | 9 | | 8 | bar | two | large | 7 | 9 | ## Step 2: Select columns indexes and values Next, determine the: - `index` \- keys to group by on the pivot table index - `columns` \- keys to group by on the pivot table column - `values` \- columns used for aggregation data of the pivot table - `aggfunc` \- functions or list of functions used for aggregation We will demonstrate a pivot table using all of the above. ## Step 3: Create the pivot table Finally, create the pivot table using method: `pivot_table()` based on the following syntax: ```python pd.pivot_table(data, values=None, index=None, columns=None, aggfunc='mean') ``` In our example we will create pivot table: - using column 'D' as aggregating values - for index we will have 2 columns - 'A' and 'B' - this creates multi-index - for columns we will use column 'C' - as aggregating function we will use Python method `sum` ```python df_pivot = pd.pivot_table(df, values='D', index=['A', 'B'], columns=['C'], aggfunc=sum) df_pivot ``` The pivot table looks like: | | C | large | small | | --- | --- | ----- | ----- | | A | B | | | | bar | one | 4.0 | 5.0 | | two | 7.0 | 6.0 | | | foo | one | 4.0 | 1.0 | | two | NaN | 6.0 | | ## Step 4: Advanced Pivot options In the previous example we saw the basic usage of the pivot\_table() method with most used options. Alternatively, we may use more options with the following default values: ```python pd.pivot_table(data, values=None, index=None, columns=None, aggfunc='mean', fill_value=None, margins=False, dropna=True, margins_name='All', observed=False, sort=True) ``` Useful pivot options are: - `fill_value` \- value to replace missing values with - `dropna` \- exclude columns whose entries are all NaN - `margins` \- add all row/columns (subtotal / grand totals) - `sort` \- sort the results ## Pivot table examples ### Pivot Table with Multiple aggfunc We can use multiple aggregation functions. The functions might be different for different columns: - 'D' - `mean` - 'E' - `min` and `max` ```python df_pivot = pd.pivot_table(df, values=['D', 'E'], index=['A', 'C'], aggfunc={'D': 'mean', 'E': [min, max]}) df_pivot ``` Result: | | | D | E | | | ----- | -------- | -------- | --- | --- | | | | mean | max | min | | A | C | | | | | bar | large | 5.500000 | 9 | 6 | | small | 5.500000 | 9 | 8 | | | foo | large | 2.000000 | 5 | 4 | | small | 2.333333 | 6 | 2 | | ### Pivot table replace NaN To replace NaN values in the pivot table we can use the parameter `fill_value`. We can replace NaN values with 0 by: ```python df_pivot = pd.pivot_table(df, values='D', index=['A', 'B'], columns=['C'], aggfunc=sum, fill_value=0) df_pivot ``` Result of replace NaN values by 0 in pivot table: | | C | large | small | | --- | --- | ----- | ----- | | A | B | | | | bar | one | 4 | 5 | | two | 7 | 6 | | | foo | one | 4 | 1 | | two | 0 | 6 | | ### Pivot table remove NaN To drop columns with NaN values we can use option `dropna=True`: ```python pd.pivot_table(df, values=['D'], index=['A'], columns=['C', 'E'], aggfunc=sum, dropna=True) ``` The result is pivot table without NaN columns: ![pandas-pivot-drop-na](https://datascientyst.com/content/images/2023/02/pandas-pivot-drop-na.png) ## Summary To summarize, in this article, we've seen an example of a creation pivot table in Pandas. We've briefly discussed syntax and options. ## Resources - [pivot\_table](https://pandas.pydata.org/docs/reference/api/pandas.pivot%5Ftable.html?ref=datascientyst.com) - [List of Aggregation Functions(aggfunc) for GroupBy in Pandas](https://datascientyst.com/list-aggregation-functions-aggfunc-groupby-pandas/) ### 334-pivot URL: https://datascientyst.com/334-pivot/ Last updated: 2023-02-15T07:19:10.000Z Pivot ### How to Fix: ValueError: Trailing Data - Pandas and JSON URL: https://datascientyst.com/fix-valueerror-trailing-data-pandas-and-json/ Last updated: 2023-02-12T08:35:18.000Z In this tutorial, we'll see how to solve a common Pandas error – `ValueError: Trailing data`. We get this error from the Pandas `read_json()` method when we try to load a JSON or JSON lines file. To fix `ValueError: Trailing data` we can try: **(1) Add parameter - `lines=True`** ```python pd.read_json('data.json', lines=True) ``` **(2) Evaluate the file line by line** ```python with open("data.json") as f: text = f.readlines() data = [eval(line) for line in text] df = pd.DataFrame(data) ``` **(3) Convert JSONl to JSON with jq** ```bash jq -s '.' data.json > out.json ``` Image below summarize the errors and some of the fixes: ![](https://datascientyst.com/content/images/2023/02/fix-valueerror-trailing-data-pandas-and-json.webp) ## 1\. Reasons - ValueError: Trailing data In Pandas and Python the error `ValueError: Trailing data` suggests that the data we are trying to load into a DataFrame is not properly formatted JSON data. There are a few common reasons why this error may occur. ### JSON lines If we try to read JSON lines file as normal JSON file without using `lines=True`: Example JSON file: ``` {"message": "Too Many Requests", "error": 429} {"message": "Too Many Requests", "error": 429} ``` ### characters outside the JSON data If there are any characters outside of the JSON data, they will cause: > ValueError: Trailing data error. Example JSON file: ``` {"message": "Too Many Requests", "error": 429}2 {"message": "Too Many Requests", "error": 429} ``` ### Inconsistent or incorrectly JSON data If JSON data is not properly formatted with correct syntax, including: - quotes - single or double quotes - values - commas separating elements Data should be consistent using only double or single quotes. Examples: ``` { "message": "Too Many Requests", "error": 429 } { "message": "Too Many Requests", "error": 429 } ``` In this example data is not in the JSON array `([])` and quotes are missing. ## 2\. Solve ValueError: Trailing data - JSON lines Depending on the case we can apply different solutions for the error. For example loading JSON lines file can be solved by adding `lines=True`: ```python import pandas as pd pd.read_json('data.json', lines=True) ``` This will solve the error and load the file: ``` {"message": "Too Many Requests", "error": 429} {"message": "Too Many Requests", "error": 429} ``` as DataFrame: | | message | error | | - | ----------------- | ----- | | 0 | Too Many Requests | 429 | | 1 | Too Many Requests | 429 | ## 3\. ValueError: Trailing data - detect errors In order to detect problematic JSON records or lines we can use the following code: ```python import pandas as pd with open('data/data_1.json') as f: content = f.readlines() data = [eval(c) for c in content] df = pd.DataFrame(data) df ``` if we try to load the JSON content of: ``` {"message": "Too Many Requests", "error": 429}2 {"message": "Too Many Requests", "error": 429} ``` We will get the following error: ``` {"message": "Too Many Requests", "error": 429}2 ^ SyntaxError: invalid syntax ``` So we can extract all problematic records and fix them. To skip problematic values check the next section. ## 4\. Handle JSON errors To skip errors in a JSON file we can read the file line by line. We can parse each line and append only good ones. For a JSON lines file with 3 rows and one of them is broken: ``` {"message": "Unknown Error", "error": 501} {"message"3: "Unknown Error", "error": 502} {"message": "Unknown Error", "error": 503} ``` We can use the following code in order to read the JSON file and skip problematic rows by: ```python import pandas as pd with open('data/data_1.json') as f: json_data = f.readlines() for row in json_data: try: data = json.loads(row) except Exception as e: pass data ``` This reads the corrupted JSON file into a DataFrame: | | message | error | | - | ------------- | ----- | | 0 | Unknown Error | 501 | | 1 | Unknown Error | 503 | As we can see line: ``` {"message"3: "Unknown Error", "error": 502} ``` Is not present in the final DataFrame. ## 5\. ValueError: Trailing data - more fixes You can also try to solve the errors also by using the following parameters: ```python pd.read_json('data.json', orient='records') pd.read_json('data.json', orient='split') pd.read_json('Data.json', encoding = 'utf-8-sig') ``` This might be helpful if you face more errors after fixing the original one: - `ValueError: Expected object or value` - `error: json.decoder.JSONDecodeError: Extra data: line 1 column 112 (char 10)` ## Conclusion To sum up, this article shows how using **proper parameters for `read_json()` method can solve the "ValueError: Trailing data" error**. We covered multiple examples and solutions for the error. If you have an interesting case or problem which is not solved by this article - please share it in the comments section below. Thanks! ### Create Count Column by value_counts in Pandas DataFrame URL: https://datascientyst.com/create-count-column-value_counts-in-pandas-dataframe/ Last updated: 2023-02-10T11:07:25.000Z In this short guide, I'll show you how to create a new count column based on `value_counts` from another column in Pandas DataFrame. There are multiple ways to count values and add them as new column: **(1) value\_counts and map** ```python counts = df['col1'].value_counts() df['col_count'] = df['col1'].map(counts) ``` **(2) group by and transform** ```python df['col_count'] = df.groupby(['col1'])['col1'].transform('count') ``` You can also read the tricky related topic: [How to Group by multiple columns, count and map in Pandas ](https://datascientyst.com/group-by-multiple-columns-count-and-map-in-pandas/). In addition we will answer on these questions: - How do I count values in a new column in pandas? - How do I create a new column based on another column value in pandas? - How do I count values in one column based on another column? Let's discuss the advantages and disadvantages of both of them in a few examples. ![](https://datascientyst.com/content/images/2023/02/create-count-column-from-value_counts-in-pandas-dataframe.png) ## Setup Let's create a sample DataFrame to count values in it's columns: ```python import pandas as pd data = {'col1': ['a', 'c', 'a', 'b', 'a', 'c'], 'col2': ['x', 'y', 'z', 'x', 'x', 'y']} df = pd.DataFrame(data) ``` DataFrame looks like: | | col1 | col2 | | - | ---- | ---- | | 0 | a | a | | 1 | c | a | | 2 | a | c | | 3 | b | e | | 4 | a | d | | 5 | c | b | ## value\_counts and map to column I prefer to use value\_counts and then map the counts to a given column. Finally we assign the values to new column: ```python counts = df['col1'].value_counts() df['col_count'] = df['col1'].map(counts) ``` we can write the same in a single line: ```python df['col_count'] = df['col1'].map(df['col1'].value_counts()) ``` result: | | col1 | col2 | col\_count | | - | ---- | ---- | ---------- | | 0 | a | a | 3 | | 1 | c | a | 2 | | 2 | a | c | 3 | | 3 | b | e | 1 | | 4 | a | d | 3 | | 5 | c | b | 2 | The advantage of this way is that we can map to different column and it's easier to read. How does it work? - the method `value_counts` calculates the count of unique values in the column `col1`. - next `map` function maps the values in `col1` to the corresponding count in the resulting Series. - the result is then assigned to a new column `col_count` in the DataFrame. ## Count and map to another column We can count values in column `col1` but map the values to column `col2`. ```python counts = df['col1'].value_counts() df['col_count'] = df['col2'].map(counts) ``` This time count is mapped to `col2` but the count is based on `col1`. This is very useful when we work with child-parent relationship: | | col1 | col2 | col\_count | | - | ---- | ---- | ---------- | | 0 | a | a | 3.0 | | 1 | c | a | 3.0 | | 2 | a | c | 2.0 | | 3 | b | e | NaN | | 4 | a | d | NaN | | 5 | c | b | 1.0 | ## group by and transform In this section we will discuss how to add a counter per group in Pandas. We can group by one column, count and then transform the results to new column: ```python df['col_count'] = df.groupby(['col1'])['col1'].transform('count') ``` We get the same result as before: | | col1 | col2 | col\_count | | - | ---- | ---- | ---------- | | 0 | a | a | 3 | | 1 | c | a | 2 | | 2 | a | c | 3 | | 3 | b | e | 1 | | 4 | a | d | 3 | | 5 | c | b | 2 | ## Performance comparison There's no difference in mid size DataFrames for both approaches: - `df.groupby(['col1'])['col1'].transform('count')` - 1.12 ms ± 15.2 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) - `df['col1'].value_counts();df['col_count'] = df['col1'].map(counts)` - 1.15 ms ± 35.6 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) Tests were done with 6000 rows. ## Conclusion In this article, we saw how to use Pandas `groupby` and `value_counts` to add a new count column in DataFrame. We also discussed how to count one column and map counts on another. ### How to Group by multiple columns, count and map in Pandas URL: https://datascientyst.com/group-by-multiple-columns-count-and-map-in-pandas/ Last updated: 2023-03-17T23:30:39.000Z To group by two or multiple columns, count unique combinations and map the result we can chain two Pandas methods: - `groupby()` - `size()` ```python df.groupby(['col1', 'col2']).size() ``` The picture below shows all the steps and the final result: ![](https://datascientyst.com/content/images/2023/02/group-by-multiple-columns-count-and-map-in-pandas.png) Let's create a sample DataFrame and explain all the steps in details: ```python import pandas as pd data = {'col1': ['a', 'c', 'a', 'b', 'a', 'c'], 'col2': ['d', 'a', 'c', 'e', 'd', 'a']} df = pd.DataFrame(data) ``` data looks like: | | col1 | col2 | | - | ---- | ---- | | 0 | a | d | | 1 | c | a | | 2 | a | c | | 3 | b | e | | 4 | a | d | | 5 | c | a | ## Group by two columns and count To group by multiple columns in Pandas and count the combinations we can chain methods: ```python df_g = df.groupby(['col1', 'col2']).size().reset_index(name='counts') ``` This gives us a new DataFrame with counts of unique combinations from the columns. We have the original columns plus new columns - `counts` which contains the occurrences: | | col1 | col2 | counts | | - | ---- | ---- | ------ | | 0 | a | c | 1 | | 1 | a | d | 2 | | 2 | b | e | 1 | | 3 | c | a | 2 | ## Map count to new column in first DataFrame Finally we can map the count to the original DataFrame. We will use method `merge` and map on two columns `['col1', 'col2']`: ```python pd.merge(df, df_g, on=['col1', 'col2']) ``` This gives us: | | col1 | col2 | counts | | - | ---- | ---- | ------ | | 0 | a | d | 2 | | 1 | a | d | 2 | | 2 | c | a | 2 | | 3 | c | a | 2 | | 4 | a | c | 1 | | 5 | b | e | 1 | Or if we like to preserve the order of the original DataFrame we can use left join - `how="left"`: ```python pd.merge(df, df_g, on=['col1', 'col2'], how='left') ``` How does it work? - we group by two columns `col1` and `col2` - the size `method` counts the number of occurrences of each combination - `reset_index` method reset the index - rename of the new column to counts and assign counts - finally map the two DataFrames on multiple columns ### Convert API Response to Pandas Dataframe - Python URL: https://datascientyst.com/convert-api-response-to-pandas-dataframe-python/ Last updated: 2023-02-09T07:36:21.000Z In this post, we will learn how to convert an API response to a Pandas DataFrame using the Python requests module. First we will read the API response to a data structure as: - CSV - JSON - XML - list of dictionaries and then we use the: - `pd.DataFrame` constructor - `pd.DataFrame.from_dict(data)` etc to create a DataFrame from that data structure. Or simply use `df=pd.read_json(url)` to convert the API to Pandas DataFrame. The image below shows the steps from API to DataFrame: ![](https://datascientyst.com/content/images/2023/02/convert-api-response-to-pandas-dataframe-python.webp) Here's simple example using a JSON weather API data: ```python import requests import pandas as pd url = "https://archive-api.open-meteo.com/v1/era5?latitude=52.52" +\ "&longitude=13.41&start_date=2021-01-01&end_date=2021-12-31&hourly=temperature_2m" response = requests.get(url) data = response.json() pd.DataFrame(data) ``` This will return the API response as Pandas DataFrame(some column are truncated for readiness): | | latitude | longitude | generationtime\_ms | utc\_offset\_seconds | timezone | | --------------- | -------- | --------- | ------------------ | -------------------- | -------- | | time | 52.5 | 13.400009 | 0.561953 | 0 | GMT | | temperature\_2m | 52.5 | 13.400009 | 0.561953 | 0 | GMT | Next we will cover the process step by step. ## Select suitable API There are hundreds of free available APIs. Below you can find several curated lists to experiment with: - [A collective list of free APIs for use in software and web development](https://github.com/public-apis/public-apis?ref=datascientyst.com) \- github collection - [What are your favorite free public API Free ones](https://www.reddit.com/r/learnpython/comments/bfz8l2/what%5Fare%5Fyour%5Ffavorite%5Ffree%5Fpublic%5Fapifree%5Fones/?ref=datascientyst.com) \- reddit post Choose one of them and test conversion from API to Pandas DataFrame with Python ## Read API in Python To read API requests in Python we will use the `requests` library. To make a GET request to an API endpoint and retrieve the response we do: ```python import requests url = "https://www.reddit.com/r/Python/comments/21q40a/why_is_pandas_so_hard/.json" response = requests.get(url) if response.status_code == 200: data = response.json() print(data) else: print("Request failed with status code:", response.status_code) ``` The `requests.get` method sends a GET request to the API at the target url. The response is stored in the response variable. We check the status code of the response by `response.status_code` attribute. If the status code is 200, the response is considered as successful. The data is converted to a JSON object using `response.json()`. If you face error like `Too Many Requests - 429` check this section ## Convert API Response to DataFrame At this step we will convert the API response to Pandas DataFrame. The conversion depends on the API response type and structure. Let's cover the case for this URL: [Sample weather API response](https://archive-api.open-meteo.com/v1/era5?latitude=52.52&longitude=13.41&start%5Fdate=2021-12-30&end%5Fdate=2021-12-31&hourly=temperature%5F2m&ref=datascientyst.com) which contains: ```json { "latitude": 52.5, "longitude": 13.400009, "generationtime_ms": 0.4019737243652344, "utc_offset_seconds": 0, "timezone": "GMT", "timezone_abbreviation": "GMT", "elevation": 38, "hourly_units": { "time": "iso8601", "temperature_2m": "°C" }, "hourly": { "time": [ "2021-01-01T00:00", "2021-01-01T01:00", "2021-01-01T02:00", "2021-01-01T03:00", ``` We can convert this JSON to DataFrame by: ```python response = requests.get(url) data = response.json() df = pd.DataFrame(data) ``` So we will have a DataFrame of two rows and multiple columns. | | latitude | longitude | generationtime\_ms | utc\_offset\_seconds | timezone | | --------------- | -------- | --------- | ------------------ | -------------------- | -------- | | time | 52.5 | 13.400009 | 0.561953 | 0 | GMT | | temperature\_2m | 52.5 | 13.400009 | 0.561953 | 0 | GMT | ## Expand nested data and plot What if you need to expand nested API data and plot it to a time series plot. As we saw there is hourly data for the temperature which is nested(you can also check the image at the start). To expand nested data with Pandas we can use method `pd.Series` \- this will convert nested data to Pandas Series: ```json pd.DataFrame(data)['hourly'].apply(pd.Series) ``` Expanded data will be accessible as new DataFrame with multiple columns: | | 0 | 1 | 2 | 3 | 4 | | --------------- | ---------------- | ---------------- | ---------------- | ---------------- | ---------------- | | time | 2021-01-01T00:00 | 2021-01-01T01:00 | 2021-01-01T02:00 | 2021-01-01T03:00 | 2021-01-01T04:00 | | temperature\_2m | 1.1 | 0.9 | 0.7 | 0.6 | 0.6 | On this page you can find more examples on expanding JSON data: [expand nested JSON or Dict in Pandas](https://datascientyst.com/normalize-json-dict-new-columns-pandas/) Finally we can plot data with: ```python pd.DataFrame(data)['hourly'].apply(pd.Series).T.set_index('time').plot() ``` ## Too Many Requests - 429 If we get error like: > {'message': 'Too Many Requests', 'error': 429} Then we can add headers to the request in order to solve it. The simplest solution is by adding user agent: ```python import requests import pandas as pd url = "https://www.reddit.com/r/Python/comments/21q40a/why_is_pandas_so_hard/.json" headers= {'User-agent': 'Mozilla/5.0 (Windows NT 6.1; Win64; x64; rv:47.0) Gecko/20100101 Firefox/47.3'} response = requests.get(url, headers = headers) data = response.json() data ``` Not using headers will return error code 429 - Too Many Requests. Using headers we are able to read the data. **Note:** Do you know that adding .json to the end of URL - converts Reddit posts into API? ## Easily read API with read\_json We can also use the method `read_json` and pass API URL to it. So just with one line of code: `df=pd.read_json(url)` Pandas can easily read and convert most API-s to Pandas DataFrame. ```python import pandas as pd url = "https://archive-api.open-meteo.com/v1/era5?latitude=52.52" +\ "&longitude=13.41&start_date=2021-01-01&end_date=2021-12-31&hourly=temperature_2m" df=pd.read_json(url) df.head() ``` The result is the same as before. We can read more about `read_json` parameters on this link: [pandas.read\_json](http://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read%5Fjson.html?ref=datascientyst.com). The most important ones are: - `path_or_buf` \- a valid JSON str, path object or file-like object - `orient` \- Indication of expected JSON string format - split - records - values You can find many useful examples on this link: [How to Read JSON Files in Pandas](https://datascientyst.com/read-json-files-pandas/) ## Conclusion We covered the most basic ways to read and convert API strings and streams to a Pandas DataFrame. ### Pandas cannot merge a series without a name URL: https://datascientyst.com/pandas-cannot-merge-a-series-without-a-name/ Last updated: 2023-02-08T14:47:56.000Z In this short guide, I'll show you how to solve Pandas error: > ValueError: Cannot merge a Series without a name The error appears when we try to merge two Pandas series without a name. To solve the error we can set name to the series either by: **(1) Rename series for the merge** ```python pd.merge(s1.rename('old'), s2.rename('new'), left_index=True, right_index=True) ``` **(2) Set name for Series** ```python s1.name = 'Series 1' ``` ## ValueError: Cannot merge a Series without a name To reproduce the error we will create two Pandas Series and try to merge them: ```python import pandas as pd s1 = pd.Series([1, 2, 3]) s2 = pd.Series([4, 5, 6]) result = pd.merge(s1, s2, left_index=True, right_index=True) ``` This results into Pandas error: ``` ValueError: Cannot merge a Series without a name ``` ## Fix - ValueError: Cannot merge a Series without a name To fix this error we can apply two solutions: ### Fix set new name Set permanent name for the Pandas series: ```python import pandas as pd s1 = pd.Series([1, 2, 3]) s1.name = 'old' s2 = pd.Series([4, 5, 6]) s2.name = 'new' result = pd.merge(s1, s2, left_index=True, right_index=True) result ``` This will change the Series: ``` 0 1 1 2 2 3 Name: old, dtype: int64 ``` The result is new DataFrame: | | old | new | | - | --- | --- | | 0 | 1 | 4 | | 1 | 2 | 5 | | 2 | 3 | 6 | ### Temporary rename series We can solve the error without changing the series by using method `rename()`: ```python import pandas as pd s1 = pd.Series([1, 2, 3]) s2 = pd.Series([4, 5, 6]) result = pd.merge(s1.rename('old'), s2.rename('new'), left_index=True, right_index=True) result ``` At the end the series is the same: ``` 0 1 1 2 2 3 dtype: int64 ``` ### How to Count Na(NaN) and non Na Values in Pandas? URL: https://datascientyst.com/count-na-nan-and-non-na-values-in-pandas/ Last updated: 2023-02-06T15:03:21.000Z In this article, we will cover how to **count NaN and non-NaN values in Pandas DataFrame or column**. Missing values in Pandas are represented by `NaN` \- not a number but sometimes are referred as: - NA - None - null We will see how to count all of them. Here is how to count NaN and non NAN values in Pandas: **(1) Count NA Values in Pandas DataFrame** ```python df.count() ``` **(2) Count non NA Values in DataFrame** ```python df.isna().sum() ``` **(3) Count NA Values in Pandas column** ```python df['col1'].count() ``` **(4) Count non NA Values in DataFrame** ```python df['col1'].isna().sum() ``` ![](https://datascientyst.com/content/images/2023/02/count-na-nan-and-non-na-values-in-pandas.webp) ## Count Na Values To count the number of NaN values in a Pandas DataFrame or Series, we can - use the `.isna()` method - then sum the resulting Boolean values( 1 = True, 0 = False): ### DataFrame To count Na values in the whole Pandas DataFrame we can apply `isna()` on every column: ```python import pandas as pd df = pd.DataFrame({'col1': ['a', None, 3, None, 5], 'col2': [None, 7, 'b', 3, 4]}) na_count = df.isna().sum() print(na_count) ``` result: ``` col1 2 col2 1 dtype: int64 ``` ### Column To count Na values in Pandas column we can sum Na values in the column: ```python df['col1'].isna().sum() ``` The result is the number of the Na values in this column - 2. ## Count non Na Values To count the number of non-NaN values in a Pandas DataFrame or Series, we can use methods: - `pandas.DataFrame.count` - `pandas.Series.count` ### DataFrame Method `count` return number of the non Na values for the whole DataFrame: ```python df.count() ``` result: ``` col1 3 col2 4 dtype: int64 ``` Not that this will count the non Na values column wise. For row-wise refer to the next section. ### Row-wise We can count non Na values in a given Pandas DataFrame row-wise by using parameter `axis=1` and pass it to `count` method: ```python df.count(axis=1) ``` result is non Na values in each row: ``` 0 1 1 1 2 2 3 1 4 2 dtype: int64 ``` ### Column To count non NaN values in Pandas column we can use the Series count method: ```python df['col1'].count() ``` as output we get the number of non Na values in col1: 3. The code above is equivalent to: ```python df['col1'].notna().sum() ``` ## Count non Na values - describe() We can use Pandas method describe to count non Na values in the whole DataFrame or column by: ```python df['col1'].describe() ``` result: ``` count 3 unique 3 top a freq 1 Name: col1, dtype: object ``` ## Count Na values - value\_counts() We can check the number of Na or non Na values also by using the method: `value_counts()`. To do so we need to pass parameter `dropna=False`: ```python df['col1'].value_counts(dropna=False) ``` result: ``` None 2 a 1 3 1 5 1 Name: col1, dtype: int64 ``` ## Count percent of missing values To count the percent of the missing values in each column of Pandas DataFrame we can use: - `isna()` - chain method `mean()` ```python df.isna().mean() ``` This will give us the percent of the Na values in the selected columns: ``` col1 0.4 col2 0.2 dtype: float64 ``` Multiply by 100 to get value between 0 and 100: ```python df.isna().mean() * 100 ``` result: ``` col1 40.0 col2 20.0 dtype: float64 ``` ## Count Na and non Na values in column To count both Na and non Na values in Pandas column we can use `isna` in combination with `value_coutns()` method: ```python df['col1'].isna().value_counts() ``` The results is number of Na and non Na values in this column: ``` False 3 True 2 Name: col1, dtype: int64 ``` ## Conclusion In this article we covered how to count the number of NaN and non NaN values in Pandas DataFrame. We saw how to count row and column-wise. We count Na values for the whole DataFrame or a single column. Finally we saw how to calculate the percent of missing values and count Na / non Na values in a column. ### How to validate IP address in Pandas URL: https://datascientyst.com/how-to-validate-ip-address-in-pandas/ Last updated: 2023-02-06T13:56:16.000Z To **validate IP addresses in a Pandas DataFrame**, we can use - the \`pd.Series.apply() method - custom function or regex Here are the 2 ways to validate IP addresses in Pandas: **(1) validate with regex** ```python df['ip'].str.contains(r"^\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}$") ``` **(2) custom validation function** ```python def validate_ip(ip): try: parts = ip.split('.') return len(parts) == 4 and all(0 <= int(part) < 256 for part in parts) except ValueError: return False except (AttributeError, TypeError): return False df['valid_ip'] = df['ip'].apply(validate_ip) ``` Suppose we work with custom DataFrame like: ```python import pandas as pd data = {'ip': ['192.168.0.1', '192.256.0.1', '192.168.0.2']} df = pd.DataFrame(data) ``` ## validate with regex To validate IP addresses with regex we have freedom of how strict the validation will be: - strict regex for IP validation - `"^(([0-9]|[1-9][0-9]|1[0-9]{2}|2[0-4][0-9]|25[0-5])\.){3}([0-9]|[1-9][0-9]|1[0-9]{2}|2[0-4][0-9]|25[0-5])$"` - basic IP validation - `r"^\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}$"` So the Pandas validation will be applied by method `str.contains` and passing the regex: ```python df['ip'].str.contains(r"^\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}$") ``` So it the basic generation we get: ``` 0 True 1 True 2 True Name: ip, dtype: bool ``` Using the strict validation we get the correct result: ```python regex = "^(([0-9]|[1-9][0-9]|1[0-9]{2}|2[0-4][0-9]|25[0-5])\.){3}([0-9]|[1-9][0-9]|1[0-9]{2}|2[0-4][0-9]|25[0-5])$" df['ip'].str.contains(regex) ``` result: ``` 0 True 1 False 2 True Name: ip, dtype: bool ``` ## validate with custom function Alternatively we can use a custom function to validate IP addresses in Pandas. This creates a new column 'valid\_ip' in the DataFrame with a Boolean value. The column indicates whether each IP address is valid or not: ```python def validate_ip(ip): try: parts = ip.split('.') return len(parts) == 4 and all(0 <= int(part) < 256 for part in parts) except ValueError: return False except (AttributeError, TypeError): return False df['valid_ip'] = df['ip'].apply(validate_ip) ``` | | ip | valid\_ip | | - | ----------- | --------- | | 0 | 192.168.0.1 | True | | 1 | 192.256.0.1 | False | | 2 | 192.168.0.2 | True | ### How to Handle Exceptions With ast.literal_eval in a Pandas URL: https://datascientyst.com/how-to-handle-exceptions-with-ast-literal_eval-in-a-pandas/ Last updated: 2023-02-06T13:28:59.000Z To handle exceptions and use `ast.literal_eval` in Pandas we can define new function: ```python import ast import pandas as pd def parse_eval(value): try: return ast.literal_eval(value) except (ValueError, SyntaxError): return value df = pd.DataFrame({'col': ['1', '2', '{"a": 3}', '[[']}) df['col'].apply(parse_eval) ``` result: ``` 0 1 1 2 2 {'a': 3} 3 [[ Name: col, dtype: object ``` Otherwise error will be raised: ```python df['col'].apply(ast.literal_eval) ``` The error is: `SyntaxError: unexpected EOF while parsing` ## ast.literal\_eval + exception To use `ast.literal_eval` in Pandas we need the `.apply()` method. This will apply the `ast.literal_eval` function to each value of a column in a Pandas DataFrame. To handle exceptions in Python we use the `try-except` block - to surround the `ast.literal_eval` call. Finally we return a default value in the except block in case of error. In the example above, the parse\_eval function takes a string value and: - returns parsed value from using `ast.literal_eval` \- when possible - returns the original string - if an exception occurs during the call to `ast.literal_eval` The `apply` method applies the `parse_eval` function to each element of the 'col' column in the DataFrame df. ### Pandas Datetime Cheat Sheet URL: https://datascientyst.com/pandas-datetime-cheat-sheet/ Last updated: 2023-02-28T18:24:49.000Z Cheat sheet for working with datetime, dates and time in Pandas and Python. The cheat sheet try to show most popular operations in a short form. There is also a visual representation of the cheat sheet. Pandas is a powerful library for working with datetime data in Python. Pandas offer variaty of attributes, methods, classes to work with date and time. The picture below illustrates most of them: ![pandas_datetime_cheat_sheet-1](https://datascientyst.com/content/images/2023/01/pandas_datetime_cheat_sheet-1.webp) The idea of this post is to save your time and effort of googling the same commands over and over. If you have an idea how to improve it - please let us know! Thank you! ## Pandas Datetime Cheat Sheet ## import datetime related libraries `from datetime import datetime, timedelta` manipulating dates and times `import pandas as pd` working with datetime Series `impot time` various time-related functions ## now & today [get current date and time in Pandas](https://datascientyst.com/get-todays-date-pandas/) `pd.to_datetime('today')` get current date and time in local timezone `pd.to_datetime('now')` current timestamp in UTC `pd.to_datetime('today').normalize()` today's date midnight `datetime.today().strftime('%d/%m/%Y')` get current date in a given format `time.strftime('%d/%m/%Y')` use module time for current date & time `datetime.datetime.now().isoformat()` get local timestamp ISO format `datetime.datetime.utcnow().isoformat()` get UTC time in ISO format ## parse date & time [parse strings to datetime](https://datascientyst.com/convert-string-datetime-timestamp-time-pandas/) `pd.to_datetime('2023-01-11 16:11:26.862697')` parse sting to datetime `pd.to_datetime(df['date'])` parse column to datetime `pd.to_datetime(df['date'], dayfirst=True)` specify parse order - True 10/11/12 is parsed as 2012-11-10 `pd.to_datetime(df['date'], yearfirst=True)` True - 10/11/12 is parsed as 2010-11-12 `pd.to_datetime(df['date'], infer_datetime_format=True)` infer the date format based on the first non-NaN item `pd.to_datetime(df['date'], format='%Y-%m-%d %H:%M:%S')` specify parsing format %Y-%m-%d %H:%M:%S `pd.to_datetime(df[['year', 'month', 'day']])` parse multiple columns ## attributes [datetime components and attributes](https://datascientyst.com/extract-year-week-datetime-pandas/) `dt.year` get year from datetime `dt.month` get month from datetime `dt.day` get day from datetime `dt.hour` get hour from datetime `dt.minute` get minute from datetime `dt.second` get second from datetime `dt.day_of_year` get day of the year from datetime `dt.dayofweek` return the day of the week `dt.is_leap_year` boolean indicator if the date belongs to a leap year `dt.daysinmonth` the number of days in the month `dt.daysinmonth` the number of days in the month ## methods popular datetime methods `dto = pd.to_datetime('2023-01-11 16:11:26.862697') dto.day_name()` return the day names with specified locale `dto.month_name()` return the month names with specified locale `dto.tz_localize('UTC')` localize tz-naive Datetime to tz-aware Datetime `dto.tz_convert('US/Central')` convert tz-aware Datetime Array/Index from one time zone to another `idx = pd.date_range('2023-01-01', periods=3) idx.to_period()` converts Datetime to Period `dto.round('H')` round operation on the data to the specified freq `dto.floor('1min')` rounds down to the nearest value `dto.ceil('1min')` rounds up to the nearest value ## calculations datetime calculation - extraction and addition of dates and time `t1 = pd.to_datetime('1/1/2023 01:00') t2 = pd.to_datetime('today') (t2 - t1).components` get time difference and return components(day, hour, minute, second) `(t2 - t1).seconds` get only seconds component `(t2 - t1).total_seconds()` get total seconds between two dates `df['date1'].dt.year - df['date2'].dt.year` calculate difference of the years `df['date'] + pd.Timedelta(days=1)` add 1 day to datetime column `df['date'] + pd.DateOffset(hours=16)` add 16 hours to datetime column `pd.Timedelta(5, 'H')` get timedelta - 5 hours `pd.Timedelta(1, 'd').total_seconds()` get 1 day delta as seconds ## select date & time [filter rows by date, year, period or time](https://datascientyst.com/filter-by-date-pandas-dataframe/) `df.set_index(['date']) df.sort_index(inplace=True, ascending=True))` set date column as index and sort for consistence `df.index = df['date']` set date as index and keep the column `df.loc['2023']` select rows by index dates in 2023 year `df.loc['2022-7']` select rows by index dates in July 2022 `df.loc['2022-1-1']` select rows by index dates by a given day `df.loc['2019' : '2022']` select all rows between two years (inclusive) `df.loc[df['date'] > '2022-01-01']` select all rows for datetime column `df.between_time('11:10','12:15')` select between start and end time `df[df['date'].between('2022', '2023')]` select rows by date column between two dates `df.loc['2022-7-1 11:10' : '2023-1-1 12:15' ]` locate by timestamps `pd.Interval(t1, t2).length` get interval length between two timestamps `pd.Interval(t1, t2).overlaps(pd.Interval(t3, t4))` check if two periods overlaps ## time zone [timezone aware timestamps and conversion](https://datascientyst.com/timezone-and-naive-timestamp-in-pandas/) `pd.Timestamp.now()` naive local time `pd.Timestamp.utcnow()` timezone aware (UTC) `pd.Timestamp.now(tz='Europe/Rome').tz_localize(None)` naive local time `pd.Timestamp.now(tz='Europe/Rome')` timezone aware local time `pd.Timestamp.utcnow().tz_localize(None) ` remove the timezone information but converting to UTC `pd.Timestamp.utcnow().tz_convert(None)` removes the timezone information resulting in naive local `dto.dt.tz_localize('+0100')` localize using offset ## read\_csv read\_csv - parsing dates `pd.read_csv('test.csv', parse_dates = ['end_date'])` try parsing columns each as a separate date column `pd.read_csv('test.csv', parse_dates = [['yy', 'mm', 'dd']])` combine columns yy,mm,dd and parse as a single date column `pd.read_csv('test.csv', parse_dates = {'date1': ['yy', 'mm', 'dd']})` parse columns yy, mm, dd as date and call result ‘date1’ `pd.read_csv('test.csv', dayfirst = True)` DD/MM format dates, international and European format `pd.read_csv('test.csv', keep_date_col = True)` if parse\_dates then keep the original columns ## date & time format Data and time directives for formatting `from datetime import datetime now = datetime.now() now.strftime('%X')` 14:56:48 || get current time and format it `%a` Thu || Abbreviated weekday (Sun) `%A` Thursday || Weekday (Sunday) `%b` Jan || Abbreviated month name (Jan) `%B` january || Month name (January) `%c` Thu Jan 5 14:50:48 2023 || Date and time `%d` 5 || Day (leading zeros) (01 to 31) `%H` 14 || 24 hour (leading zeros) (00 to 23) `%I` 2 || 12 hour (leading zeros) (01 to 12) `%j` 5 || Day of year (001 to 366) `%m` 1 || Month (01 to 12) `%M` 53 || Minute (00 to 59) `%p` PM || AM or PM `%S` 48 || Second (00 to 29) `%U` 1 || Week number (00 to 53) `%w` 4 || Weekday (0 - Sun to 6 - Sat) `%W` 1 || Week number (00 to 53) `%x` 01/05/23 || Date `%X` 14:56:48 || Time `%y` 23 || Year without century (00 to 99) `%Y` 2023 || Year (2008) `%Z` GMT || Time zone (GMT) `%%` % || A literal % character ## Working with datetime Here are some **common tasks you might want to do with datetime data in Pandas**: 1. **Create a datetime column in a Pandas DataFrame** \- we can use the `pd.to_datetime` method to convert a column of strings to a column of datetime objects. For example: ```python import pandas as pd df = pd.DataFrame({'date_string': ['2022-01-01', '2022-01-02', '2022-01-03']}) df['date'] = pd.to_datetime(df['date_string']) ``` result is a datetime Series and new column: ``` 0 2022-01-01 1 2022-01-02 2 2022-01-03 Name: date_string, dtype: datetime64[ns] ``` 1. **Extract the year, month, or day from a datetime column** \- we can use the `dt` attribute of a datetime columns to extract the year, month, or day as a separate column. For example: ```python df['year'] = df['date'].dt.year df['month'] = df['date'].dt.month df['day'] = df['date'].dt.day ``` 1. **Filter a DataFrame by a datetime column**: `dt` attribute can be used to filter a DataFrame by a specific date range. For example: ```python # filter rows where the date year is 2023 df_2023 = df[df['date'].dt.year == 2023] # filter rows where the date is in January 2023 df_january_2023 = df[(df['date'].dt.year == 2023) & (df['date'].dt.month == 1)] ``` 1. **Aggregate data by a datetime column** \- we can use the `groupby()` method to aggregate data by a datetime column. Example: ```python # calculate the mean value of a column for each year df.groupby(df['date'].dt.year)['value'].mean() # calculate the sum of a column for each month df.groupby(df['date'].dt.month)['value'].sum() ``` ![](https://datascientyst.com/content/images/2023/01/pandas_datetime_cheat_sheet.svg) ### How to Create a Bag of Words in Pandas Python URL: https://datascientyst.com/create-a-bag-of-words-pandas-python/ Last updated: 2023-01-10T21:25:11.000Z In this short guide, I'll show you how to create a bag of words with Pandas and Python. You can find a example of bag of words using the `sklearn` library: ```python from sklearn.feature_extraction.text import CountVectorizer import pandas as pd text = ['The fox jumps over the lazy dog.', 'Dog and fox are lazy!'] data = {'text': text} df = pd.DataFrame(data) vectorizer = CountVectorizer() bow = vectorizer.fit_transform(df['text']) print(bow) count_array = bow.toarray() features = vectorizer.get_feature_names() df = pd.DataFrame(data=count_array, columns=features) ``` Below you can find the result of the code: ![bag_of_words_python](https://datascientyst.com/content/images/2023/01/bag_of_words_python.webp) In the next steps I'll explain the process in more detail. ## What is a Bag of Words? **A bag of words is a way to represent text data in tabular form as numerical features**. You can also find a quick solution only with Pandas and Python below: ```python import pandas as pd from collections import Counter text = ['Periods of rain', 'Mostly cloudy, a little rain', 'Mostly cloudy', 'Intervals of clouds and sun', 'Sunshine and mild'] df = pd.DataFrame({'text': text}) pd.DataFrame(df['text'].str.split().apply(Counter).to_list()) ``` The input DataFrame is: | | text | | - | ---------------------------- | | 0 | Periods of rain | | 1 | Mostly cloudy, a little rain | | 2 | Mostly cloudy | | 3 | Intervals of clouds and sun | | 4 | Sunshine and mild | The output bag of words represented again as DataFrame: | | Periods | of | rain | Mostly | cloudy, | a | little | cloudy | Intervals | clouds | and | sun | Sunshine | mild | | - | ------- | --- | ---- | ------ | ------- | --- | ------ | ------ | --------- | ------ | --- | --- | -------- | ---- | | 0 | 1.0 | 1.0 | 1.0 | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | | 1 | NaN | NaN | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | NaN | NaN | NaN | NaN | NaN | NaN | NaN | | 2 | NaN | NaN | NaN | 1.0 | NaN | NaN | NaN | 1.0 | NaN | NaN | NaN | NaN | NaN | NaN | | 3 | NaN | 1.0 | NaN | NaN | NaN | NaN | NaN | NaN | 1.0 | 1.0 | 1.0 | 1.0 | NaN | NaN | | 4 | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | 1.0 | NaN | 1.0 | 1.0 | **Note:** For a bag of words data preprocessing is needed. Remove special characters and stopwords, convert words to lowercase, stemming etc. The diagram below shows how to create a bag of words from multiple documents. Each item in the list is considered as separate document. ![](https://datascientyst.com/content/images/2023/01/create-a-bag-of-words-pandas-python.webp) ## Setup First lets create a sample DataFrame for this example: ```python from sklearn.feature_extraction.text import CountVectorizer import pandas as pd text = ['Periods of rain', 'Mostly cloudy, a little rain', 'Mostly cloudy', 'Intervals of clouds and sun', 'Sunshine and mild'] data = {'text': text} df = pd.DataFrame(data) ``` result: | | text | | - | ---------------------------- | | 0 | Periods of rain | | 1 | Mostly cloudy, a little rain | | 2 | Mostly cloudy | | 3 | Intervals of clouds and sun | | 4 | Sunshine and mild | We are going to import `CountVectorizer` from `sklearn`. ## Step 1: Initialize the vectorizer Next we are going to initialize the vectorizer of `sklearn`: ```python vectorizer = CountVectorizer() ``` ### lowercase At this step we can do customization like on/off of `lowercase` conversion: ```python vectorizer = CountVectorizer(lowercase=False) ``` ### Stop words or adding **custom stop words** for the bag of words: ```python vectorizer = CountVectorizer(stop_words= ['on', 'off']) ``` even using stop words based on languages: ```python coun_vect = CountVectorizer(stop_words='english') ``` ### max\_df / min\_df The abbreviation `df` in `max_df` / `min_df` stands for `document frequency`. - `max_df = 0.75` \- ignore terms that appear in more than 75% of the documents. - `max_df = 0.5` \- ignore terms that appear in less than 50% documents. When using a float in the range `[0.0, 1.0]` they refer to the document frequency. They can be used also in sense of `max_df = 10` \- which means ignore terms that appear in less than 10 documents ```python coun_vect = CountVectorizer(max_df=1) ``` ## Step 2: Fit and transform the text data Next step is to fit and transform the text data to create a bag of words: ```python bow = vectorizer.fit_transform(df['text']) ``` This creates a bag of words from the DataFrame column like: ``` (0, 8) 1 (0, 7) 1 (0, 9) 1 (1, 9) 1 (1, 6) 1 (1, 2) 1 (1, 4) 1 (2, 6) 1 (2, 2) 1 (3, 7) 1 (3, 3) 1 (3, 1) 1 ``` This is a sparse matrix, where: - each row represents a document - each column represents a word - the values in the matrix represent the number of times that word appears in that document. So in `(0, 8) 1` we have: - 0 is the number of the document - first document - 8 number of the feature - periods - 1 - is the count - 1 occurrence ## Step 3: Get features names and counts We can get the word list from the vectorizer, by calling the method `get_feature_names()`: ```python # get count array and features count_array = bow.toarray() features = vectorizer.get_feature_names() # create DataFrame as bag of words df = pd.DataFrame(data=count_array, columns=features) ``` This will create a DataFrame where: - each column is a word - row represents the documents - values are number of times each word is present | and | clouds | cloudy | intervals | little | mild | mostly | of | periods | rain | sun | sunshine | | --- | ------ | ------ | --------- | ------ | ---- | ------ | -- | ------- | ---- | --- | -------- | | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | | 0 | 0 | 1 | 0 | 1 | 0 | 1 | 0 | 0 | 1 | 0 | 0 | | 0 | 0 | 1 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | | 1 | 1 | 0 | 1 | 0 | 0 | 0 | 1 | 0 | 0 | 1 | 0 | | 1 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 1 | ## Normalize the bag of words Alternatively, we can also use `TfidfVectorizer` from `scikit-learn` that creates a bag of words by: - first counting the frequency of each word in each document - then normalizing the resulting counts by dividing by the total number of words in the document. The resulting values are the term frequency-inverse document frequency (TF-IDF) values: ```python from sklearn.feature_extraction.text import TfidfVectorizer import pandas as pd text = ['small dog', 'cute cat', 'cute dog', 'cat'] data = {'Text':text} df = pd.DataFrame(data) vectorizer = TfidfVectorizer() bow = vectorizer.fit_transform(df['Text']) print(bow) ``` result: ``` (0, 2) 0.6191302964899972 (0, 3) 0.7852882757103967 (1, 0) 0.7071067811865475 (1, 1) 0.7071067811865475 (2, 1) 0.7071067811865475 (2, 2) 0.7071067811865475 (3, 0) 1.0 ``` Getting bag of words as a DataFrame with normalized values: ```python count_array = bow.toarray() features = vectorizer.get_feature_names() df = pd.DataFrame(data=count_array, columns=features) ``` | cat | cute | dog | small | | -------- | -------- | -------- | -------- | | 0.000000 | 0.000000 | 0.619130 | 0.785288 | | 0.707107 | 0.707107 | 0.000000 | 0.000000 | | 0.000000 | 0.707107 | 0.707107 | 0.000000 | | 1.000000 | 0.000000 | 0.000000 | 0.000000 | ## Conclusion To summarize, in this article, we've seen examples of bags of words. We've briefly covered what a bag of words is and how to create it with Python, Pandas and scikit-learn. And finally, we've seen how to normalize a bag of words with scikit-learn. ### Instantly Turn Web Pages into Beautiful Dashboards with Python URL: https://datascientyst.com/instantly-turn-web-pages-into-beautiful-dashboards-python/ Last updated: 2024-04-03T09:33:47.000Z ## Intro I recently had the need to monitor multiple web pages and filter information from them. This is a rather simple task but it's time consuming and error prone. Every time I do it, it takes time to find the right data, analyze it and save it. In addition, often I would like to monitor multiple web sources simultaneously without distraction in a homogeneous style. I did research for this problem but nothing was close enough to my needs. What I need is to build a beautiful dashboard from multiple web sites. In the past, I was using CRON jobs, Python scripts and Jupyter notebooks to collect data in one place. Finally I found a better solution which extracts data from web pages and turns them into a dashboard. You will also learn how to turn any Jupyter Notebook into a dashboard in seconds. ![](https://datascientyst.com/content/images/2022/12/turn-web-pages-into-beautiful-dashboards-python-1.webp) ## Setup ![turn-web-pages-into-beautiful-dashboards-python](https://datascientyst.com/content/images/2022/12/turn-web-pages-into-beautiful-dashboards-python.webp) ## Prerequisite ### Voilà [Voilà](https://pypi.org/project/voila/?ref=datascientyst.com) is a Python package which turns Jupyter notebooks into standalone web applications. It can be used as a standalone app with the new Jupyter kernel or inside Jupyter. It can be installed by: ```bash pip install voila ``` Voilà provides a JupyterLab extension that displays a Voilà preview of your Notebook in a side-pane. To install the extension from source, run the following command. ```bash jupyter labextension install @voila-dashboards/jupyterlab-preview ``` ### voila-gridstack [voila-gridstack](https://pypi.org/project/voila-gridstack/?ref=datascientyst.com) is gridstack-based template for Voilà. ```bash pip install voila-gridstack ``` ## 1\. Scraping with Pandas Our goal is to make quick and easy extracting data from multiple sources into a single dashboard. There are many ways to scrape data with Python. The simplest and easiest way to scrape tabular data is by using Pandas. We will cover two different options: - basic extraction - adding user agent ### Pandas scrape tables Pandas offers handy method [pandas.read\_html](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fhtml.html?ref=datascientyst.com) which reads HTML tables into a list of DataFrames. By default extracts all tables from a given URL: ```python import pandas as pd url_cur = 'https://en.wikipedia.org/wiki/List_of_countries_by_forest_area' df_ls = pd.read_html(url_cur) df_ls[0] ``` | | Region | 1990 | 2000 | 2010 | 2020 | | - | --------------------------------- | ------- | ------- | ------- | ------- | | 0 | World | 4236433 | 4158050 | 4106317 | 4058931 | | 1 | Europe (including Russia) | 994319 | 1002268 | 1013982 | 1017461 | | 2 | South America | 973666 | 922645 | 870154 | 844186 | | 3 | North America and Central America | 755279 | 752349 | 754190 | 752710 | | 4 | Africa | 742801 | 710049 | 676015 | 636639 | | 5 | Asia | 585393 | 587410 | 610960 | 622687 | | 6 | Oceania | 184974 | 183328 | 181015 | 185248 | ### Pandas read\_html + user agent Some websites with cause Pandas method `read_html()` to return: > HTTPError: HTTP Error 403: Forbidden In order to solve this problem we will add user agent and use package `requests`: ```python import pandas as pd import requests url_cur = 'https://en.wikipedia.org/wiki/List_of_countries_by_forest_area' header = { "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/50.0.2661.75 Safari/537.36", "X-Requested-With": "XMLHttpRequest" } r = requests.get(url_cur, headers=header) ls_cur = pd.read_html(r.text) ``` Adding a user agent as headers solves the `HTTPError: HTTP Error 403: Forbidden` returned from Pandas. ## 2\. Styling and plotting Python and Pandas offer multiple ways for styling and creating beautiful visualizations. For simplicity we will mention only two in this article. ### DataFrame as heatmap The first approach is to use method \`.style.background\_gradient() that can be used create nice looking heatmaps: ```python df.style.background_gradient(cmap='Greens', subset=str_cols)\ .background_gradient(cmap='Blues', subset='Price') ``` To find more about it check out: [How to Display Pandas DataFrame As a Heatmap ](https://datascientyst.com/display-pandas-dataframe-heatmap/) ### Quick visualization with seaborn The second way for creating quick and nice visualizations is by using libraries like `seaborn`. The advantage of `seaborn` is simplicity of usage and diversity of plots. To learn more about different visualization options and styles refer to: [Pandas Visualization Cheat Sheet](https://datascientyst.com/pandas-visualization-cheat-sheet/) ## 3\. Turn Jupyter Notebook into Dashboard Once we have data collected and all visualizations are ready for use - we can start building our dashboard. 0:00 /0:15 1× ### Open voila-gridstack editor To open voila-gridstack editor we have two options in JupyterLab. The first way is by: - right clicking the notebook - Open with - Voilà Gridstack Alternatively we can use a button on the right top of an opened notebook. ### drag-and-drop cells We can see the notebook and the editor side by side. We can select a cell and move it to the voila-gridstack editor. Then we can resize or move the cell in the voila-gridstack editor. The image below shows the process. ### Open new voila window Once we are happy with the outlook of the dashboard we can save it. Finally we can open it as a separate window or use it as a standalone application. ## Conclusion There are many options for building a dashboard within the Python ecosystem. Voila offers a quick and easy way to render Jupyter notebooks as a dashboard . In my experience, Voila is the best choice for beginners and people with medium experience. I hope this article will be a useful guide for people interested in building their own dashboards for their practical problems. Feel free to leave a comment to ask a question. ## Resources - [Voila - Rendering of live Jupyter notebooks with interactive widgets.](https://github.com/voila-dashboards/voila?ref=datascientyst.com) - [voila-gridstack - A gridstack-based template for voila-gridstack](https://github.com/voila-dashboards/voila-gridstack?ref=datascientyst.com) ### 👋 Getting Started with Data Science Project URL: https://datascientyst.com/getting-started-with-data-science-project/ Last updated: 2023-01-25T12:09:31.000Z Are you one of those people who think that learning is difficult? One of the hardest parts about learning a new skill is getting started. Data science can be intimidating and scary at first. Yes, it's true: there are many things to learn - statistics, mathematics, programming - this can be overwhelming. Based on my experience, the best is to start with a small project to get the taste of it. One small step at a time, a few small steps each day, and we can learn complex topics over time. We all define success in a different way. For me being successful is to do what I love to do. Yet, we are all encouraged by completing successful tasks or projects. To complete data science project you need to perform 5 important steps: ![data_science_project_steps](https://datascientyst.com/content/images/2022/12/data_science_project_steps.webp) - Idea & concept - Understand data - Solve the problem - Present it - Keep it simple, stupid! Does it seem simple enough? No! I know, because I've been there. Knowing is not enough, we need to practice. Only by practice can confidence and expertness be assured. To start, pick one topic, project or idea, and just start it. What could be better than a fun visualization that can be shared with friends? A simple project to gain confidence and boost you learning process: Good example of successful presentation is the following reddit example: > I haven't seen GOT myself but I love how according to viewers and ratings they built a beautiful airplane and then flew it right into a mountain wall. [Game of Thrones - user ratings s8 - worst, e9 best](https://www.reddit.com/r/dataisbeautiful/comments/zeu725/oc%5Fgame%5Fof%5Fthrones%5Fuser%5Fratings%5Fs8%5Fworst%5Fe9%5Fbest/?ref=datascientyst.com) One of the most desired requests for learning materials we see at Data Scientyst is about data science projects. Let us know what else you would like to see more from us by: - [Google form for ideas](https://docs.google.com/forms/d/e/1FAIpQLSc7rHn8Dh-Zu8twMDDFcYrLrK1kps2NlPuZcEHVcYAYWhXCMA/viewform?ref=datascientyst.com) - [Contact Us](mailto:datascientyst@gmail.com) ### 434-newsletter URL: https://datascientyst.com/newsletter/ Last updated: 2022-12-21T15:20:50.000Z Newsletter ### Pandas Visualization Cheat Sheet URL: https://datascientyst.com/pandas-visualization-cheat-sheet/ Last updated: 2022-11-25T11:43:27.000Z **This visualization cheat sheet is a great resource to explore data visualizations with Python, Pandas and Matplotlib**. The Python ecosystem provides many packages for producing high-quality plots, graphs and visualizations. In this guide, we will discuss the basics and a few popular visualization choices. The article starts with the basic steps for creating visualization. Next these steps are covered in detail. The end of this article has useful resources for visualizations - free books, guides, galleries. There are summary images showing multiple visualizations at once. The goal of this guide is to help you building and customizing data visualizations. Let's dive into visualization cheat sheet. Below you can find most popular plots from Seaborn: ![seaborn_barplot_lineplot_heatmap](https://datascientyst.com/content/images/2022/11/seaborn_barplot_lineplot_heatmap.webp) ## How to create good visualization Python offers a ton of options and ways to visualize and summarize data which makes Python a natural choice for Data science. Every great story starts with an idea. The same is with the visualization - we need idea and steps to follow to create great visualization. 1. Idea 2. Collect and select data 3. Data cleaning 4. Prepare data 1. dimensions 2. X and Y axis data 3. plot type - boxplot, line chart 5. Select tool 6. Select style and color palette 7. Customize the plot 1. title 2. labels 3. data format 4. size Let your data and plots tell your story. ## Data Setup In this post we will use two DataFrames: - DataFrame with random numbers - Seaborn dataset Creating DataFrame with 1000 numbers using normal distribution: ```python import pandas as pd import numpy as np ts = pd.Series(np.random.randn(1000), index=pd.date_range("1/1/2000", periods=1000)) df = pd.DataFrame(np.random.randn(1000, 4), index=ts.index, columns=list("ABCD")) df = df.head(5) ``` result: | | A | B | C | D | | ---------- | ---------- | -------- | ---------- | ---------- | | 2000-01-01 | \-0.004858 | 0.618783 | \-0.960541 | \-0.118617 | | 2000-01-02 | \-0.476119 | 0.972206 | 0.457535 | \-0.099867 | | 2000-01-03 | \-0.043310 | 0.218806 | \-0.751540 | \-0.501480 | | 2000-01-04 | \-1.913368 | 0.143043 | 1.140921 | \-0.569990 | | 2000-01-05 | 1.076793 | 0.809909 | 1.009482 | 0.716194 | Seaborn DataFrame ```python import seaborn as sns glue = sns.load_dataset("glue").pivot("Model", "Task", "Score") df_tit = sns.load_dataset("titanic") penguins = sns.load_dataset("penguins") df_sns = sns.load_dataset('flights') ``` data looks like is: | | year | month | passengers | | - | ---- | ----- | ---------- | | 0 | 1949 | Jan | 112 | | 1 | 1949 | Feb | 118 | | 2 | 1949 | Mar | 132 | | 3 | 1949 | Apr | 129 | | 4 | 1949 | May | 121 | ## Pandas visualization cheat sheet Pandas can visualize DataFrame by using the method `plot()`. It has a backend specified by the option `plotting.backend` \- by default - `matplotlib`. Documentation for this method is available on this link: [DataFrame.plot](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.plot.html?ref=datascientyst.com). ### Setup, import, save We need several imports to plot data with Python, Pandas and Matplotlib. ```python import pandas as pd import matplotlib.pyplot as plt ``` Save and show figure: ```python plt.savefig('plot.png') plt.savefig('plot.png', transparent=True) #transparent plt.show() ``` ### Figure To create new figure in Matplotlib with a given size: - set size in inches - figaspect will determine the width and height for a figure that would fit array preserving aspect ratio ```python fig = plt.figure() fig = plt.figure(figsize=(10,5)) # size in inches fig = plt.figure(figsize=plt.figaspect(3.0)) w, h = figaspect(2.) fig = Figure(figsize=(w,h)) ``` ### Axes To add and delete axes ```python fig.add_axes() fig.add_axes(ax) fig.delaxes(ax) ``` ### Subplot Working with Subplots in Matplotlib ```python add_subplot(nrows, ncols, index, **kwargs) add_subplot(pos, **kwargs) add_subplot(ax) add_subplot() ax1 = fig.add_subplot(111) #row/col/ix ax2 = fig.add_subplot(112) fig, axes = plt.subplots(nrows=2,ncols=2) fig, axes = plt.subplots(nrows=4) ``` ### Matplotlib Markers ```python ax.scatter(x,y,marker= ".") ax.plot(x,y,marker= "o") ``` Available markers in Matplotlib: - `"."` \- point - `"o"` \- circle - `"v"` \- triangle down - `"s"` \- square - `"D"` \- diamond - `"*"` \- star marker To find more markers we can visit: [matplotlib.markers API](https://matplotlib.org/stable/api/markers%5Fapi.html?ref=datascientyst.com) ### Linestyles in Matplotlib To find different line styles we can visit: [set\_linestyle](https://matplotlib.org/stable/api/%5Fas%5Fgen/matplotlib.lines.Line2D.html?ref=datascientyst.com#matplotlib.lines.Line2D.set%5Flinestyle): - `'-'` \- solid line - `'--'` \- dashed line - `'-.'` \- dash-dotted line - `':'` \- dotted line ```python x = df['A'] y = df['B'] plt.plot(x,y,linewidth=5.0) plt.plot(x,y,linestyle= 'solid' , color='y') plt.plot(y,x,ls= '--') plt.plot(y,x,'--' ,x**2,y**3,'-.' ) ``` Plot different line styles with Matplotlib: - color - style ![linestyles-in-matplotlib](https://datascientyst.com/content/images/2022/11/linestyles-in-matplotlib.webp) ```python from math import * import numpy as np x = np.arange(0,1.0,0.01) y1 = np.sin(2*pi*x) y2 = np.sin(4*pi*x) lines = plt.plot(x, y1, x, y2) plt.setp(lines, linewidth=2) ``` Plot different color lines with `plt.setp`: ![matplotlib-linestyles](https://datascientyst.com/content/images/2022/11/matplotlib-linestyles.webp) ```python import matplotlib.pyplot as plt import numpy as np import pandas as pd cycler = plt.cycler(linestyle=['-', ':', '--', '-.'], color=['r', 'b', 'y', 'g']) fig, ax = plt.subplots() ax.set_prop_cycle(cycler) df.plot(ax=ax) plt.show() ``` multiple line styles and colors: ![matplotlib-multiple-linestyles-colors](https://datascientyst.com/content/images/2022/11/matplotlib-multiple-linestyles-colors.webp) ### Pandas plot Series To plot Pandas Series we can call method `plot()` directly on the Series: ```python import pandas as pd s = pd.Series([5, 7, 2, 4, 1]) ax = s.plot(kind='bar', figsize=(10,5)) ``` The result is bar plot from the Series: ![pandas_plot_series](https://datascientyst.com/content/images/2022/11/pandas_plot_series.webp) ### Pandas plot DataFrame We can Plot DataFrame in Pandas by calling the method `plot()`. ```python ax = df.plot() ``` By default all numeric columns will be used for the visualization: ![pandas-plot-dataframe](https://datascientyst.com/content/images/2022/11/pandas-plot-dataframe.webp) Prior plotting DataFrame we can: - select only the columns that will be plot - set index or select X and Y data - format and clean data To plot DataFrame as bar plot, using the year and month as X axis with custom figure size we can do: ```python df.set_index(['year', 'month']).plot(kind='bar', figsize=(30,10)) ``` This will plot number of passengers as Y axis: ![pandas_plot_multiple_columns_x_axis](https://datascientyst.com/content/images/2022/11/pandas_plot_multiple_columns_x_axis.webp) ### Title, Labels, Legend ```python ax.set_xlabel('Year') ax.set_ylabel('Passenger') ax.set_title('Passengers per year') ax.legend(labels, loc='best') ax.set(title= 'Title', ylabel= 'Y label', xlabel= 'X axis') ``` ### Ticks ```python import pandas as pd s = pd.Series([5, 7, 2, 4, 1]) ax = s.plot(figsize=(10,5)) ax.yaxis.set(ticks=range(1,9,3), ticklabels=['min', 'mid', 'max']) ax.tick_params(axis= 'y', direction= 'out', length=5) ``` result: ![pandas_plot_ticks](https://datascientyst.com/content/images/2022/11/pandas_plot_ticks.png) ### Margins, Limits ```python import pandas as pd s = pd.Series([5, 7, 2, 4, 1]) ax = s.plot(figsize=(10,5)) ax.margins(x=0.5,y=0.5) # ax.axis('equal') # Equal axis size # ax.set_xlim(1,5) # set x limit ax.set(xlim=[-1,5],ylim=[1,9]) # set x & y limits ``` result: ![pandas_plot_margin_limit](https://datascientyst.com/content/images/2022/11/pandas_plot_margin_limit.webp) ### Parameters - `x` \- X axis data - `y` \- Y axis data - `kind='bar'` \- plot type - `ax` \- - `figsize` \- plot size in inches - `subplots=True` \- subplots for each column - `sharex=False` \- in case of subplots - should X axis be shared - `layout=(3,2)` \- shape of the subplots - number of rows and columns - `title` \- title of the plot - `xticks/yticks` \- values to use for the xticks/yticks - `xlabel` / `ylabel` \- name to use for the labels on x-axis/y-axis - `fontsize=12` \- font size for title - `color='green'` \- plot color - `color = ['lightblue', 'r', 'y']` more parameters on: [plot()](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.plot.html?ref=datascientyst.com) Example for parameter `subplots=True`: ![pandas_subplots_True](https://datascientyst.com/content/images/2022/11/pandas_subplots_True.webp) ### Display two plots - side by side ```python from matplotlib import pyplot as plt # First plot ax = plt.subplot() plt.pie( data=df, x='A') plt.title( 'bar' ) plt.show() # Second plot ax = plt.subplot() plt.scatter( data=df, x='A', y='B' ) plt.title( 'scatter' ) plt.show() ``` Two plots side by side - pie chart and scatter plot: ![display-two-plots-side-by-side](https://datascientyst.com/content/images/2022/11/display-two-plots-side-by-side.webp) ### Matplotlib subplots ```python import matplotlib.pyplot as plt import numpy as np data = np.array([1, 4, 2, 3, 2]) plt.subplot(121) plt.plot(data) data = np.array([5, 7, 3, 8, 3]) plt.subplot(122) plt.plot(data) plt.show() ``` result: ![matplotlib_subplots](https://datascientyst.com/content/images/2022/11/matplotlib_subplots.webp) ### Grids of Subplots ```python for i in range(1, 5): plt.subplot(2, 2, i) plt.text(0.5, 0.5, str((2, 2, i)), font size=18, ha='center') ``` result: ![matplotlib_grid_subplots](https://datascientyst.com/content/images/2022/11/matplotlib_grid_subplots.webp) ### Types of Pandas plots List of the available chart types for `plot()` method. The plot type can be set as parameter - `kind`: - `line` : line plot (default) - `bar` : vertical bar plot - `barh` : horizontal bar plot - `hist` : histogram - `box` : boxplot - `kde` : Kernel Density Estimation plot - `density` : same as 'kde' - `area` : area plot - `pie` : pie plot - `scatter` : scatter plot (DataFrame only) - `hexbin` : hexbin plot (DataFrame only) Every available plot in Pandas is shown below: ![pandas-plot-dataframe_all_types](https://datascientyst.com/content/images/2022/11/pandas-plot-dataframe_all_types_bar_line_histogram.webp) The code below generates all plots: ```python import pandas as pd import numpy as np import matplotlib.pyplot as plt import seaborn as sns; sns.set_theme() ts = pd.Series(np.random.randn(1000), index=pd.date_range("1/1/2000", periods=1000)) df = pd.DataFrame(np.random.randn(1000, 4), index=ts.index, columns=list("ABCD")) df = df.head(5) plots = [ 'line', 'hist', 'box', 'kde', 'density', 'area', 'pie', 'scatter', 'barh', 'bar', 'hexbin'] cols = df.columns row_num = 3 col_num = 4 row_n = -1 col_n = 0 fig, axes = plt.subplots(row_num, col_num, squeeze=False, figsize=(20,14)) for ix, plot in enumerate(plots): axes[row_n, col_n].title.set_size(20) axes[row_n, col_n].title.set_color('red') col_n = ix % col_num if col_n == 0: row_n = row_n + 1 if plot not in ['area', 'pie', 'scatter', 'hexbin']: df.plot(kind=plot, ax=axes[row_n, col_n], title=plot, figsize=(30,12)) elif plot == 'area': df.plot(kind=plot, ax=axes[row_n, col_n], title=plot, stacked=False) elif plot == 'pie': series = pd.Series(3 * np.random.rand(4), index=["a", "b", "c", "d"], name="series") series.plot.pie(ax=axes[row_n, col_n], title=plot); elif plot == 'scatter': df.plot(kind=plot, ax=axes[row_n, col_n], title=plot, x=['A'], y=['B']) elif plot == 'hexbin': df.plot(kind=plot, ax=axes[row_n, col_n], title=plot, x=['A'], y=['B']) plt.show() ``` ## Seaborn vs Matplotlib Seaborn is based on Matplotlib. It enhances Matplotlib by simplifying the plot process and adding new features. On the image below we can see all Seaborn plots like: - [heatmap](https://seaborn.pydata.org/generated/seaborn.heatmap.html?ref=datascientyst.com) - [boxplot](https://seaborn.pydata.org/generated/seaborn.boxplot.html?ref=datascientyst.com) - [barplot](https://seaborn.pydata.org/generated/seaborn.barplot.html?ref=datascientyst.com) - [lineplot](https://seaborn.pydata.org/generated/seaborn.lineplot.html?ref=datascientyst.com) - [histogram](https://seaborn.pydata.org/generated/seaborn.histplot.html?ref=datascientyst.com) - pie chart - currently Seaborn doesn't have pie chart ![seaborn_barplot_lineplot_heatmap](https://datascientyst.com/content/images/2022/11/seaborn_barplot_lineplot_heatmap.webp) ### Seaborn setup We can import and load datasets with seaborn by next code: ```python import seaborn as sns glue = sns.load_dataset("glue").pivot("Model", "Task", "Score") df_sns = sns.load_dataset('flights') df_tit = sns.load_dataset("titanic") penguins = sns.load_dataset("penguins") ``` ### Seaborn heatmap To plot heatmap with seaborn we can do simply: ```python sns.heatmap(glue) ``` ![seaborn_heatmap](https://datascientyst.com/content/images/2022/11/seaborn_heatmap.webp) ### Seaborn boxplot Plotting boxplot in Seaborn is as easy as: ```python df_tit = sns.load_dataset("titanic") sns.boxplot(x=df["age"]) ``` ### Seaborn barplot For barplot we need to: - select data source - X and Y axis - data ```python penguins = sns.load_dataset("penguins") sns.barplot(data=penguins, x="island", y="body_mass_g") ``` ### Seaborn histogram Seaborn histogram is called - `histplot`: ```python penguins = sns.load_dataset("penguins") sns.histplot(data=penguins, x="flipper_length_mm") ``` ### Seaborn multiple plots To plot multiple visualization in Seaborn side by side we can do: ```python import pandas as pd import seaborn as sns import matplotlib.pyplot as plt df = sns.load_dataset('penguins') sns.scatterplot(data=df, x='bill_length_mm', y='bill_depth_mm', hue='sex') plt.show() sns.scatterplot(data=df, x='flipper_length_mm', y='body_mass_g', hue='sex') plt.show() ``` result: ![seaborn_multiple_plots](https://datascientyst.com/content/images/2022/11/seaborn_multiple_plots.webp) ## Colors Python comes with a huge variety of named colors and palettes like the one shown below. To find the full list of colors check: - [Full List of Named Colors in Pandas and Python](https://datascientyst.com/full-list-named-colors-pandas-python-matplotlib/) - [How to Get a List of N Different Colors and Names in Python/Pandas](https://datascientyst.com/get-list-of-n-different-colors-names-python-pandas/) Changing colors in matplotlib: ```python color = 'red' mpl.rcParams['text.color'] = color mpl.rcParams['axes.labelcolor'] = color mpl.rcParams['xtick.color'] = 'y' mpl.rcParams['ytick.color'] = 'b' ``` ![matplotlib-colors](https://datascientyst.com/content/images/2022/11/matplotlib-colors.webp) ## Data Visualization Libraries - [Matplotlib ](https://matplotlib.org/?ref=datascientyst.com) \- most popular and widely-used plotting library - [Examples](https://matplotlib.org/stable/plot%5Ftypes/index.html?ref=datascientyst.com) - [Reference](https://matplotlib.org/stable/index.html?ref=datascientyst.com) - [Cheat Sheets](https://matplotlib.org/cheatsheets/?ref=datascientyst.com) - [seaborn ](https://seaborn.pydata.org/?ref=datascientyst.com) \- based on Matplotlib. High-level interface for drawing attractive and informative statistical graphics - [Gallery](https://seaborn.pydata.org/examples/index.html?ref=datascientyst.com) - [Tutorial](https://seaborn.pydata.org/tutorial.html?ref=datascientyst.com) - [plotly](https://plotly.com/?ref=datascientyst.com) \- Plotly's Python graphing library makes interactive, publication-quality graphs. - [Getting Started](https://plotly.com/python/getting-started?ref=datascientyst.com) - [Examples](https://plotly.com/python/?ref=datascientyst.com) - [bokeh](https://bokeh.org/?ref=datascientyst.com) \- interactive visualization library - [Gallery](https://docs.bokeh.org/en/latest/docs/gallery.html?ref=datascientyst.com) - [First steps](https://docs.bokeh.org/en/latest/docs/first%5Fsteps.html?ref=datascientyst.com) - [Vega-Altair](https://altair-viz.github.io/?ref=datascientyst.com) \- produces beautiful and effective visualizations with a minimal amount of code. - [Example Gallery](https://altair-viz.github.io/gallery/index.html?ref=datascientyst.com) - [User Guide](https://altair-viz.github.io/user%5Fguide/data.html?ref=datascientyst.com) - [pygal](https://www.pygal.org/en/stable/?ref=datascientyst.com) is a dynamic SVG charting library written in python - [geoplotlib](https://pypi.org/project/geoplotlib/?ref=datascientyst.com) \- open-source Python library for visualizing geographical data. ## Materials/Books for data visualization ### Free Materials - [Fundamentals of Data Visualization](https://clauswilke.com/dataviz/?ref=datascientyst.com) - [github](https://github.com/clauswilke/dataviz?ref=datascientyst.com) - [an introduction to VISUALIZING DATA](https://beforebefore.net/scima300/f15/media/visualizingdata.pdf?ref=datascientyst.com) - [Getting Started with Data Visualization](https://web.stanford.edu/group/toolingup/cgi-bin/toolkit/wp-content/uploads/2011/03/GMcGhee%5Ftoolingup%5FDataVis%5F110506.pdf?ref=datascientyst.com) - [Principles of Data Visualization I ](https://indico.cern.ch/event/681081/contributions/2790760/attachments/1729504/2794629/Principles-of-Visualization-Course-Pt1-Full.pdf?ref=datascientyst.com) - [Visualization with Matplotlib - Python Data Science Handbook](https://jakevdp.github.io/PythonDataScienceHandbook/04.00-introduction-to-matplotlib.html?ref=datascientyst.com) ### Paid Books - [Storytelling with Data: A Data Visualization Guide for Business Professionals](https://www.storytellingwithdata.com/books?ref=datascientyst.com) \- Don't simply show your data—tell a story with it! - [Storytelling with Data: Let's Practice!](https://www.storytellingwithdata.com/books?ref=datascientyst.com) \- Influence action through data! - [Information is Beautiful ](https://informationisbeautiful.net/books/?ref=datascientyst.com) \- A stunning visual journey through the most amazing, beautiful, and positive things happening in the modern world. - [Better Data Visualizations A Guide for Scholars, Researchers, and Wonks](http://cup.columbia.edu/book/better-data-visualizations/9780231193115?ref=datascientyst.com) \- essential strategies to create more effective data visualizations ### Collections - [R Data Visualization Books](https://www.bigbookofr.com/data-visualization.html?ref=datascientyst.com) \- 17 books for Data Visualization and R ### Resources Data Visualization - [https://www.python-graph-gallery.com/](https://www.python-graph-gallery.com/?ref=datascientyst.com) \- collection of hundreds of charts made with Python - [https://www.data-to-viz.com/](https://www.data-to-viz.com/?ref=datascientyst.com) \- leads you to the most appropriate graph for your data - [https://www.reddit.com/r/dataisbeautiful/](https://www.reddit.com/r/dataisbeautiful/?ref=datascientyst.com) \- DataIsBeautiful is a reddit for visualizations that effectively convey information - [https://www.reddit.com/r/dataisugly/](https://www.reddit.com/r/dataisugly/?ref=datascientyst.com) \- bad data visualizations ### FIFA World Cup 2022: Data-Driven Analyze (Twitter) URL: https://datascientyst.com/data-driven-approach-to-analyze-fifa-world-cup-2022-twitter/ Last updated: 2023-03-19T09:31:53.000Z Today is the first day of the World Cup 2022 which takes place in Qatar from 20-th of November to 18-th of December. 32 teams will compete in eight groups for the prize. In this post we will use data-driven approach to analyze the teams and what people twit for the new football event. What we can learn from this post is how to extract tweets with Python library [tweepy](https://pypi.org/project/tweepy/?ref=datascientyst.com) and use Pandas to analyse them. This article will walk through how to collect meta data, scrape tweets and read all tweets into a Pandas DataFrame. Because the Twitter setup and authentication process is complex, we will not cover it in details. This is a sample data science project for confident beginners :) ## Requirements Libraries needed for this project: ``` pip install tweepy pip install python-dotenv pip install pycountry pip install emoji-country-flag ``` ## 1\. Collect meta data The first part of the process is collecting finalist countries and their flags. If you haven't used Pandas to scrape web before, use the following links: - [pandas.read\_html](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fhtml.html?ref=datascientyst.com) - [Easily extract tables from websites with Pandas and Python](https://www.youtube.com/watch?v=OXA%5FZD1gR6A&ref=datascientyst.com) ### 1.1 Collect finalists Once we find the page where we can collect data, we can use it in the next Python code: ```python from pandas.io.html import read_html page = 'https://en.wikipedia.org/wiki/2022_FIFA_World_Cup_qualification' wikitables = read_html(page) len(wikitables) ``` Method `read_html` reads HTML tables into a list of DataFrame objects. The result is 30. As we can see the page: [2022 FIFA World Cup qualification](https://en.wikipedia.org/wiki/2022%5FFIFA%5FWorld%5FCup%5Fqualification?ref=datascientyst.com) contains 30 tables. The one which contains our data is available under number 1 - we can find that by inspecting the page or displaying DataFrames: ```python df_fin = wikitables[1] df_fin.head() ``` Now we have all countries which plays on the World Cup 2022 as Pandas DataFrame: | | Team | Method ofqualification | Date ofqualification | Totaltimesqualified | Lasttimequalified | Currentconsecutiveappearances | Previous bestperformance | | - | ------- | ---------------------- | -------------------- | ------------------- | ----------------- | ----------------------------- | -------------------------------------- | | 0 | Qatar | Hosts | 2 December 2010 | 1 | – | 1 | – | | 1 | Germany | UEFA Group J winners | 11 October 2021 | 20\[a\] | 2018 | 18 | Winners (1954, 1974, 1990, 2014) | | 2 | Denmark | UEFA Group F winners | 12 October 2021 | 6 | 2018 | 2 | Quarter-finals (1998) | | 3 | Brazil | CONMEBOL winners | 11 November 2021 | 22 | 2018 | 22 | Winners (1958, 1962, 1970, 1994, 2002) | | 4 | France | UEFA Group D winners | 13 November 2021 | 16 | 2018 | 7 | Winners (1998, 2018) | ### 1.2 Get country flags Python offers elegant way to collect information about counrtries like: - flag - abbreviations - official name The first library is named: [pycountry](https://pypi.org/project/pycountry/?ref=datascientyst.com). This Python library provides the ISO databases for the standards: - Languages - Countries - Subdivisions of countries - Currencies Currently contains about 250 countries. From this library we would use fuzzy matching to collect the country details: ```python import pycountry pycountry.countries.search_fuzzy('Brazil') ``` which results into: ``` [Country(alpha_2='BR', alpha_3='BRA', flag='🇧🇷', name='Brazil', numeric='076', official_name='Federative Republic of Brazil')] ``` For our needs we will collect the flags for the 32 teams on the World Cup: ```python countries = [] for fin in finalists: try: country = pycountry.countries.search_fuzzy(fin)[0] flag = country.flag country = {fin:flag} except: print(fin) countries.append(country) print(countries) ``` Which results into: ``` England [{'Qatar': '🇶🇦'}, {'Germany': '🇩🇪'}, {'Denmark': '🇩🇰'}, {'Brazil': '🇧🇷'}, {'France': '🇫🇷'}, {'Belgium': '🇧🇪'}, {'Serbia': '🇷🇸'}, {'Spain': '🇪🇸'}, {'Croatia': '🇭🇷'}, {'Switzerland': '🇨🇭'}, {'Switzerland': '🇨🇭'}, {'Netherlands': '🇳🇱'}, {'Argentina': '🇦🇷'}, {'Iran': '🇮🇷'}, {'South Korea': '🇰🇷'}, {'Saudi Arabia': '🇸🇦'}, {'Japan': '🇯🇵'}, {'Uruguay': '🇺🇾'}, {'Ecuador': '🇪🇨'}, {'Canada': '🇨🇦'}, {'Ghana': '🇬🇭'}, {'Senegal': '🇸🇳'}, {'Poland': '🇵🇱'}, {'Portugal': '🇵🇹'}, {'Tunisia': '🇹🇳'}, {'Morocco': '🇲🇦'}, {'Cameroon': '🇨🇲'}, {'United States': '🇺🇸'}, {'Mexico': '🇲🇽'}, {'Wales': '🇦🇺'}, {'Australia': '🇦🇺'}, {'Costa Rica': '🇨🇷'}] ``` There is open issue for England: [search\_fuzzy fails for 'England'](https://github.com/flyingcircusio/pycountry/issues/126?ref=datascientyst.com) Alternatively we can use another library called: [emoji-country-flag](https://pypi.org/project/emoji-country-flag/?ref=datascientyst.com). ```python import flag flag.flag("FR") ``` result: ``` '🇫🇷' ``` Finally we need to get English flag as dict `{'England': '󠁧󠁢󠁥󠁮󠁧🏴󠁧󠁢󠁥󠁮󠁧󠁿'}` and append it to the list. ## 2\. Extract Tweets For this step we are going to use library: [tweepy](https://pypi.org/project/tweepy/?ref=datascientyst.com). Depending on the Twitter access we can use two different approaches to extract tweets: - Essential access - Twitter API v2 endpoints only - Elevated access - everything ### 2.1 Setup a Twitter Developer Account To setup new Twitter Developer Account we can use these articles as a reference: - [tweepy authentication](https://docs.tweepy.org/en/latest/authentication.html?ref=datascientyst.com#twitter-api-v2) - [How To Extract Data From The Twitter API Using Python](https://towardsdatascience.com/how-to-extract-data-from-the-twitter-api-using-python-b6fbd7129a33?ref=datascientyst.com). ### 2.2 Essential access Using this approach we need only BEARER\_TOKEN from Twitter. So we can search for tweets by query `'#FIFAWorldCup'`. To make this code work you need to replace `os.environ["BEARER_TOKEN"]` with your token. Alternatively you can create `.env` file with the following syntax - the file should be in the same folder as the script: ``` API_KEY="xxx" API_KEY_SECRET="xxx" BEARER_TOKEN="xxx" ACCESS_TOKEN="xxx" ACCESS_TOKEN_SECRET="xxx" ``` Then to extract tweets run the following code: ```python import tweepy import pandas as pd bearer_token = os.environ["BEARER_TOKEN"] client = tweepy.Client(bearer_token=bearer_token) """ More examples https://github.com/twitterdev/getting-started-with-the-twitter-api-v2-for-academic-research/blob/main/modules/5-how-to-write-search-queries.md """ query = '#FIFAWorldCup' tweets = client.search_recent_tweets(query=query, tweet_fields=['context_annotations', 'created_at'], max_results=10) pd.DataFrame(tweets.data) ``` By default the Twitter API will return max of 100 results. In order to get more than 100 tweets we are using built-in tweepy.Paginator: ```python import tweepy client = tweepy.Client(bearer_token=bearer_token) query = '#FIFAWorldCup' tweets = tweepy.Paginator(client.search_recent_tweets, query=query, tweet_fields=['context_annotations', 'created_at'], max_results=100).flatten(limit=10000) ls = [] for tweet in tweets: ls.append(tweet) import pandas as pd df = pd.DataFrame(ls) ``` The tweets stored as Pandas DataFrame: | | created\_at | id | text | | - | ------------------------- | ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | 2022-11-20 11:14:41+00:00 | 1594288115992043520 | RT @sushimmii: @rahmdess27 @BTS\_twt DREAMERS BY JUNGKOOK\\nDREAMERS STREAMING PARTY\\n\\n#Dreamers2022 #FIFAWorldCup @BTS\_twt #DreamersByJungkoo… | | 1 | 2022-11-20 11:14:41+00:00 | 1594288115895914496 | RT @FIFAWorldCup: "See you at the Opening!" - Jung Kook \\n\\nToday, 5.30pm local time. \\n\\n#Qatar2022 \| #FIFAWorldCup https://t.co/3oulTUthBV | | 2 | 2022-11-20 11:14:41+00:00 | 1594288115635539969 | @kingnyamjoon DREAMERS BY JUNGKOOK\\nDREAMERS STREAMING PARTY\\n\\n#Dreamers2022 #FIFAWorldCup @BTS\_twt #DreamersByJungkook #JungKook #정국 | | 3 | 2022-11-20 11:14:41+00:00 | 1594288115581358081 | RT @btsarmy39521947: Jungkook 💜\\nI’m so addicted with #Jungkook of \\n@BTS\_twt\\n's new single "Dreamers" for the #FIFAWorldCup ! The lyri… | | 4 | 2022-11-20 11:14:41+00:00 | 1594288115283214336 | RT @Bangtan\_twt\_com: @btssomma DREAMERS BY JUNGKOOK\\nDREAMERS STREAMING PARTY\\n\\n#Dreamers2022 #FIFAWorldCup #DreamersByJungkook #JungK… | ### 2.3 Elevated access You need to apply for Elevated access to work with the code below. Otherwise you will face an error: > Forbidden: 403 Forbidden > 453 - You currently have Essential access which includes access to Twitter API v2 endpoints only. If you need access to this endpoint, you’ll need to apply for Elevated access via the Developer Portal. You can learn more here: [https://developer.twitter.com/en/docs/twitter-api/getting-started/about-twitter-api#v2-access-leve](https://developer.twitter.com/en/docs/twitter-api/getting-started/about-twitter-api?ref=datascientyst.com#v2-access-leve) Code: ```python import os import tweepy from dotenv import load_dotenv, find_dotenv from pathlib import Path path='./conf/.env' load_dotenv(dotenv_path=path,verbose=True) consumer_key = os.environ["API_KEY"] consumer_secret = os.environ["API_KEY_SECRET"] access_token = os.environ["ACCESS_TOKEN"] access_token_secret = os.environ["ACCESS_TOKEN_SECRET"] auth = tweepy.OAuth1UserHandler( consumer_key, consumer_secret, access_token, access_token_secret ) api = tweepy.API(auth) tweets = api.search_tweets("World cup", tweet_mode="extended") for tweet in tweets: try: print(tweet.retweeted_status.full_text) print("=====") except AttributeError: print(tweet.full_text) print("=====") ``` ## 3\. Analyse Extracted Data Finally we will use data collected in previous two steps. ### 3.1\. Finding the most mentioned countries To find the most mentioned countries we can use loop over the teams and count the mentions: ```python count_team = [] for team in df_fin['Team']: count_team.append(df[df['text'].str.contains(team)].shape[0]) df_fin['count'] = count_team df_fin.sort_values('count', ascending=False).set_index('Team').tail(31)['count' ].plot(kind='bar') ``` We exlude Qatar from the chart as it has more than 3500 mentions: ![data-driven-approach-to-analyze-fifa-world-cup-2022-twitter](https://datascientyst.com/content/images/2022/11/data-driven-approach-to-analyze-fifa-world-cup-2022-twitter.webp) We can see that the other top mentioned country is Ecuador. This is expected as this is going to be the first game: Qatar - Ecuador. ### 3.2\. Searching by flags We can use flags to search in the tweets - as we know emojis are very popular in the Twitter jargon: ```python df[df.text.str.contains(flag.flag("FR"))].shape df[df.text.str.contains(flag.flag("BR"))].shape ``` So we get that Brazile is mentioned 48 times while France is mentioned 45 times. ## Conclusion There are several new ideas we learned with this post:. - how to easily extract tabular data with Pandas - several new Python libraries to make your live much easier - how to scrape tweets with few line of codes - data-driven approach to solve everyday problems Finally, I love the motto of 2022 World Cup: > "Expect Amazing" ### Style Pandas DataFrame Like a Pro (Examples) URL: https://datascientyst.com/style-pandas-dataframe-like-pro-examples/ Last updated: 2026-01-23T17:15:55.000Z In this tutorial, we'll discuss the **basics of Pandas Styling and DataFrame formatting.** We will also check frequently asked questions for DataFrame styles and formats. We'll start with basic usage, methods, parameters and then see a few Pandas styling examples. Next, we'll learn how to beautify DataFrame and communicate data more efficiently. Additionally, we'll discuss tips and also learn some advanced techniques like cell or column highlighting. Hope that you will learn invaluable tips for Pandas styling and formatting like: ![pandas-dataframe-style-heatmap-min-max-3-colors](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-heatmap-min-max-3-colors.webp) and ![](https://datascientyst.com/content/images/2022/11/style-pandas-dataframe-like-data-scientist-examples.webp) Which one is better for the last image? Without formatting or with? If you need to check formatting and styling for pivot tables in check this: [Apply Formatting and Borders to Pivot table in Pandas](https://datascientyst.com/how-to-apply-formatting-and-borders-to-pivot-table-in-pandas/) ## Setup In this tutorial we will work with the Seaborn dataset for flights. Which can be loaded with method `sns.load_dataset()` ```python import seaborn as sns import pandas as pd df = sns.load_dataset('flights') pd.pivot_table(df, index='year', columns='month') ``` We will convert the initial DataFrame to a pivot table. This will give us a better DataFrame for styling. Initial Data looks like: | | year | month | passengers | | - | ---- | ----- | ---------- | | 0 | 1949 | Jan | 112 | | 1 | 1949 | Feb | 118 | | 2 | 1949 | Mar | 132 | | 3 | 1949 | Apr | 129 | | 4 | 1949 | May | 121 | While the pivot table is - having all years like rows and all months as columns (below data is truncated): | | passengers | | | | | | | | ----- | ---------- | --- | --- | --- | --- | --- | --- | | month | Jan | Feb | Mar | Apr | May | Jun | Jul | | year | | | | | | | | | 1949 | 112 | 118 | 132 | 129 | 121 | 135 | 148 | | 1950 | 115 | 126 | 141 | 135 | 125 | 149 | 170 | | 1951 | 145 | 150 | 178 | 163 | 172 | 178 | 199 | | 1952 | 171 | 180 | 193 | 181 | 183 | 218 | 230 | ## 1\. How do I style a Pandas DataFrame? To style a Pandas DataFrame we need to use `.style` and pass styling methods. This returns a Styler object and not a DataFrame. We can control the styling by parameters and options. We can find the most common methods and parameters for styling in Pandas in the next section. The syntax for the Pandas Styling methods is: ```python df.style.highlight_null(null_color="blue") ``` ### 1.1 Combine Pandas styling methods Styling methods can be chained so we can replace NaN values and highlight them in red background at once: ```python df1.style.format(na_rep='').highlight_null(null_color="red") ``` Formatting of the last method in the chain takes action. `NaN` values with be highlighted in blue: ```python df1.style.highlight_null(null_color="red").highlight_null(null_color="blue") ``` ### 1.2 Why using Pandas styling methods Several reasons why to use Pandas styling methods: - focus attention on the important data and trends - style change only visual representation and not the data - beauty attracts attention - you will show better understanding of the subject - choosing correct styling is power data science skill ## 2\. Pandas styling methods Let's start with most popular Pandas methods for DataFrame styling like: - `format(na_rep='', precision=2)` \- general formatting of DataFrame - missing values, decimal precision - `background_gradient()` \- style DataFrame as heatmap - `highlight_min()` \- highlight min values across rows or columns - `highlight_max()` \- highlight max values across rows or columns - `bar()` \- display column as bar - `highlight_null()` \- highlight missing values in DataFrame - `set_table_styles()` \- style the entire table, columns, rows or specific HTML selectors. - `set_table_attributes()` \- set the table attributes added to the `` HTML element - `set_properties()` \- Set defined CSS-properties to each `
` HTML element for the given subset Some methods are still available by will be deprecated in future: - `set_na_rep()` - replace NaN values - replaced by Styler.format(na\_rep=..) - `set1_precision()` - set decimal precision - replaced by Styler.format(precision=..) - `hide_index()` - hide the index - Styler.hide(axis='index') ### 2.1 Pandas Styling Parameters Popular parameters for Pandas styling: - `axis` - 0 - row wise - 1 - column wise - None - DataFrame wise - `subset` - column/row names on which the styling will be applied - `cmap` - the color scheme for the styling - examples - summer, Greens, ocean - to find more options - enter wrong value and get all options from the exception - `vmin` / `vmax` \- the minimum and maximum value ### 2.2 Pandas Format DataFrame To format the text display value of DataFrame cells we can use method: `styler.format()`: ```python df.style.format(na_rep='MISS', precision=3) ``` Result is replacing missing values with string 'MISS' and set float precision to 3 decimal places: ![pandas-dataframe-style-format](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-format.webp) Another format example - add percentage to the numeric columns: ```python df.style.format("{:.2%}", subset=1) ``` ![pandas-dataframe-style-format-string](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-format-string.webp) Default values for this method: - `styler.format.formatter`: default None. - `styler.format.na_rep`: default None. - `styler.format.precision`: default 6. - `styler.format.decimal`: default “.”. - `styler.format.thousands`: default None. - `styler.format.escape`: default None. We can combine method format with lambda to format the columns: ```python .format({"col_1": lambda x:x.upper()}) ``` This will convert the column `col_1` to upper case. ### 2.3 DataFrame as heatmap To **convert Pandas DataFrame to a beautiful Heatmap** we can use method `.background_gradient()`: ```python df_p.style.background_gradient() ``` The result is colored DataFrame which show us that number of passengers grow with the increase of the years: ![pandas-dataframe-style-heatmap](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-heatmap.webp) One more example using parameters `vmin` and `vmax`: ```python df_p.style.background_gradient(cmap = "RdYlGn", vmin = 104, vmax = 622) ``` which create descriptive visual table: ![pandas-dataframe-style-heatmap-min-max-3-colors](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-heatmap-min-max-3-colors.webp) More example about: [How to Display Pandas DataFrame As a Heatmap](https://datascientyst.com/display-pandas-dataframe-heatmap/) ### 2.4 DataFrame column as bar chart To convert Pandas column to bar visualization inside the DataFrame output we can use method `bar`: ```python df.head().style.bar(subset=['passengers'], cmap='summer') ``` We can see a clear pattern by using the bar styling. Passenger increase in the summer and decrease in the winter months: ![pandas-dataframe-style-bar](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-bar.webp) ### 2.5 Highlight max values To highlight max values in Pandas DataFrame we can use the method: `highlight_max()`. By default highlights max values per column: ```python df_p.style.highlight_max() ``` To highlight max values per row we need to pass - `axis=1`. In case of max value in more than one cell - all will be highlighted: ```python df_p.style.highlight_max(axis=1) ``` The max values are highlighted in yellow. Which makes easy to digest data: ![pandas-dataframe-style-highlight_max](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-highlight_max.webp) ### 2.6 Highlight min values To highlight the min values we can use: `highlight_min()`. It's pretty similar to the max values from above. We can find the absolute minimum value by - `axis=None`: ```python df_p.style.highlight_min(axis=None) ``` This will focus the attention on the absolute min value: ![pandas-dataframe-style-highlight_min](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-highlight_min.webp) ### 2.7 Highlight NaN values in DataFrame To highlight NaN values in a Pandas DataFrame we can use the method: `.highlight_null()`. Selecting the color for the NaN highlight is available with parameter - `null_color="blue"`: To prepare the NaN values we use: ```python df1 = df.head() df1.iloc[[1,2],1] = pd.NA ``` and then we highlight them by: ```python df1.style.highlight_null(null_color="red") ``` ### 2.8 Replace NaN values in Pandas styling To replace NaN values with string in a Pandas styling we can use two methods: - `.format(na_rep='')` - `.set_na_rep()` \- this one will be deprecated in future Replacing NaN values in styling with empty spaces: ```python df1.style.format(na_rep='') ``` or: ```python df1.style.set_na_rep("") ``` Note: This method will soon be deprecated - so you can use: `Styler.format(na_rep=..)` to avoid future errors ![pandas-dataframe-style-highlight_nan](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-highlight_nan.webp) ### 2.9 Set title to DataFrame To set title to Pandas DataFrame we can use method: `set_caption()` ```python df.style.set_caption("DataScientYst 2022") ``` Title is added to the DataFrame: ![pandas-dataframe-style-title](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-title.webp) ### 2.10 Set table styles in DataFrame To set table styles and properties of Pandas DataFrame we can use method: `set_table_styles()` To apply table styles only for specific columns we can select the columns by: ```python df.style.set_table_styles({ 1: [{'selector': '', 'props': [('color', 'red')]}], 4: [{'selector': 'td', 'props': 'color: blue;'}] }) ``` Columns 1 and 4 are changed: To apply new table style and properties we can use HTML selectors like: - `*` \- select all - `th` \- header - `tr` \- row - `td` \- cell ```python styles = [{'selector':"*", 'props':[ ("font-family" , 'Mono'), ("font-size" , '15px'), ("margin" , "15px auto"), ("border" , "2px solid #ccc"), ("border-bottom" , "2px solid #00eeee")]}] df.style.set_table_styles(styles) ``` the result from both examples: ![pandas-dataframe-style-table-format](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-table-format.webp) ## 3\. Pandas apply format To apply format on Pandas DataFrame we can use methods: - `.apply()` - `.applymap()` Example for `applymap` used to color column in red: ```python def highlight_cols(s): color = 'salmon' return 'background-color: %s' % color df.style.applymap(highlight_cols, subset=pd.IndexSlice[:, [1]]) ``` Use `apply()` to format string values: ```python def highlight_strings(s): return ['background-color: salmon' if type(val) == str else 'background-color: lime' for val in s] df.style.apply(highlight_strings) ``` results are visible below: ![pandas-dataframe-style-apply-applymap](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-apply-applymap.webp) ## 4\. Pandas styling examples & FAQ ### 4.1 How do I beautify a DataFrame in Python? To beautify Pandas DataFrame we can combine different methods to create visual impact. First let's create simple DataFrame from numbers from 0 to 24: ```python import numpy as np import pandas as pd a = np.arange(25).reshape(5,5) df = pd.DataFrame(a) ``` Next we will define the function `color_divisible` \- and apply it on the DataFrame. Then we will change the table properties like - headers, rows etc: ```python def color_divisible(num, div): background = 'background-color: PaleGreen' if num % div == 0 else '' return background props = [('font-size', '12pt'),('border-style','solid'),('border-width','1px')] df.style.applymap(color_divisible, div=3)\ .set_table_attributes('style="font-size: 20px"')\ .set_table_styles([{'selector': 'th', 'props': props}]) ``` The result is attached to the image. ![pandas-beautify-dataframe-style](https://datascientyst.com/content/images/2022/11/pandas-beautify-dataframe-style.webp) Second example on - how to beautify DataFrame. Coloring the table headers, values and changing border styles: ```python styler = df.style.applymap(color_divisible, div=3)\ .set_caption("DataScientYst 2022") props = [('color', 'black'), ('border-style','solid'), ('border-width','1px')] sel_all = {'selector': '*', 'props': props} sel_th= {'selector': 'th', 'props': [('background-color', 'gold')]} styler.set_table_attributes('style="font-size: 25px"') styler.set_table_styles([sel_all, sel_th ]) styler ``` The beautified DataFrame is below: ![pandas-dataframe-style-table-properties](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-table-properties.webp) ### 4.2 How do you color a column in Pandas? Depending on the results and data we can use different techniques to color Pandas columns. We already saw(will see) how to color column: - in a single color with [applymap/apply](https://datascientyst.com/p/2bdc1422-5cdf-4e63-baa6-1010266bea0e/#3-pandas-apply-format) - as [heatmap](https://datascientyst.com/p/2bdc1422-5cdf-4e63-baa6-1010266bea0e/#23-dataframe-as-heatmap) with `.background_gradient()` and subset - as [bar](https://datascientyst.com/p/2bdc1422-5cdf-4e63-baa6-1010266bea0e/#24-dataframe-column-as-bar-chart) with `.bar(subset=['passengers'], cmap='summer')` ### 4.3 How do I change the color of a DataFrame in Python? Usually I prefer to change the color of DataFrame by using combination of: - numeric values - `.background_gradient()` and subset - `.bar(subset=['passengers'], cmap='summer')` - highlight groups in Pandas - [Pandas group by and highlight](https://datascientyst.com/pandas-dataframe-background-color-based-condition-value-alternate-row-color-based-group/) - [Color boolean values in DataFrame](https://datascientyst.com/style-boolean-values-different-colors-in-pandas/) ### 4.4 How do I highlight cells in Pandas? For conditional formatting of DataFrame I prefer to use the built-in style functions. If something is not covered as functionality - then I will use custom function with: - `apply` - `applymap` to highlight cells like we saw: - `highlight_cols` - `highlight_strings` ### 4.5 How to pretty print Pandas DataFrame To pretty print Pandas DataFrame we can use the built in function `.to_markdown()`: ```python print(df.to_markdown()) ``` result: | | 0 | 1 | 2 | 3 | 4 | | - | -- | -- | -- | -- | -- | | 0 | 0 | 1 | 2 | 3 | 4 | | 1 | 5 | 6 | 7 | 8 | 9 | | 2 | 10 | 11 | 12 | 13 | 14 | | 3 | 15 | 16 | 17 | 18 | 19 | | 4 | 20 | 21 | 22 | 23 | 24 | Or Python library [tabulate](https://pypi.org/project/tabulate/?ref=datascientyst.com): ```python from tabulate import tabulate print(tabulate(df, headers='keys', tablefmt='simple')) ``` result: ``` 0 1 2 3 4 -- --- --- --- --- --- 0 0 1 2 3 4 1 5 6 7 8 9 2 10 11 12 13 14 3 15 16 17 18 19 4 20 21 22 23 24 ``` ### 4.6 Render Pandas DataFrame as HTML To render Pandas DataFrame as HTML we can use method - `.to_html()`: ```python print(df.to_html()) ``` Then we can use the HTML table code generated from the DataFrame: ```html ``` ### 4.7 Export DataFrame format as Excel table To export DataFrame as Excel table, keep styles and formatting we can use method: `.to_excel('style.xlsx', engine='openpyxl')`: ```python styler = df.style.applymap(color_divisible, div=3) styler.to_excel('style.xlsx', engine='openpyxl') ``` The code above will create a Pandas style. Then we export the styles to a file named `style.xlsx`. ![pandas-dataframe-style-excel-export-format](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-excel-export-format.webp) ### 4.8 Format Pandas DataFrame as Excel table Excel has pre-built table formats - altering color rows. To format DataFrame as Excel table we can do: ```python css_alt_rows = 'background-color: lightgreen; color: black;' css_indexes = 'background-color: green; color: white;' df.style.set_table_styles([ {'selector': 'tr:nth-child(even)', 'props': css_alt_rows}, {'selector': 'th', 'props': css_indexes}, ]) ``` Find the results - DataFrame styled as Excel table below: ![pandas-dataframe-style-excel-format](https://datascientyst.com/content/images/2022/11/pandas-dataframe-style-excel-format.webp) ## 5\. Global Display Options in Pandas To change Pandas display option we can use several methods like: - [get\_option()](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.get%5Foption.html?ref=datascientyst.com#pandas.get%5Foption "pandas.get_option") / [set\_option()](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.set%5Foption.html?ref=datascientyst.com#pandas.set%5Foption "pandas.set_option") \- get/set the value of a single option. - [reset\_option()](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.reset%5Foption.html?ref=datascientyst.com#pandas.reset%5Foption "pandas.reset_option") \- reset one or more options to their default value. - [describe\_option()](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.describe%5Foption.html?ref=datascientyst.com#pandas.describe%5Foption "pandas.describe_option") \- print the descriptions of one or more options. - [option\_context()](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.option%5Fcontext.html?ref=datascientyst.com#pandas.option%5Fcontext "pandas.option_context") \- execute a codeblock with a set of options that revert to prior settings after execution. ```python import pandas as pd pd.options.display.max_rows # 15 pd.options.display.max_rows = 999 pd.set_option("display.max_rows", 999) ``` show more columns and rows(or [show all columns and rows in Pandas](https://datascientyst.com/pandas-show-all-columns-rows/): ```python with pd.option_context("display.max_rows", 1000, "display.max_columns", 50): print(pd.get_option("display.max_rows")) print(pd.get_option("display.max_columns")) ``` To find more for Pandas options we can refer to the official documentation: [Pandas options and settings](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/options.html?ref=datascientyst.com) ## 6\. Pandas Styling tips Finally we will cover several tips for styling Pandas DataFrames: - don't overdo it - use styles when needed. To many colors might distract the person who will digest the information - control the styles with parameters like - `subset` \- format subset of DataFrame - `axis` \- rows or columns - ask for feedback before sharing it on larger audience - add titles, legends - anything which is required for correct understanding of the styles/data - format columns based on the data - amounts - distance - temperature - dates - select colors carefully - research on other people work and share your work Share your tips as comments below the article! Thank you! ## References 1. API reference - [pandas.io.formats.style.Styler.format](https://pandas.pydata.org/docs/reference/api/pandas.io.formats.style.Styler.format.html?ref=datascientyst.com) 2. User Guide - [Table Visualization — pandas 1.5.1 documentation - PyData](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/style.html?ref=datascientyst.com) 3. User Guide - [Styling — pandas 1.1.5 documentation](https://pandas.pydata.org/pandas-docs/version/1.1/user%5Fguide/style.html?ref=datascientyst.com) 4. [Options and settings in Pandas](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/options.html?ref=datascientyst.com) 5. [display articles - DataScientyst](https://datascientyst.com/display/) 6. [Table articles - DataScientyst](https://datascientyst.com/table/) 7. [Styling articles - DataScientyst](https://datascientyst.com/styling/) ### Data Science Project for beginners in 15 minutes URL: https://datascientyst.com/dataisbeautiful-the-absolute-quality-of-breaking-bad/ Last updated: 2023-03-17T23:02:16.000Z This article shows how to scrape, analyze and visualize movie data from IMDb. We will learn how to use Python and Pandas in order to collect, transform and present data in a beautiful way. ## Objective The second goal is to follow all steps in order to create popular DataIsBeautiful visualization: - [\[OC\] The absolute quality of Better Call Saul](https://www.reddit.com/r/dataisbeautiful/comments/yfphz4/oc%5Fthe%5Fabsolute%5Fquality%5Fof%5Fbetter%5Fcall%5Fsaul/?ref=datascientyst.com) - [\[OC\] The absolute quality of Breaking Bad](https://www.reddit.com/r/dataisbeautiful/comments/fwjces/oc%5Fthe%5Fabsolute%5Fquality%5Fof%5Fbreaking%5Fbad/?ref=datascientyst.com) So we will try to create similar artwork as the one posted on Reddit: ![](https://datascientyst.com/content/images/2022/10/dataisbeautiful-the-absolute-quality-of-breaking-bad.webp) You can find notebook for this article: [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/softhints/Pandas-Exercises-Projects/blob/main/project/dataisbeautiful-the-absolute-quality-of-breaking-bad.ipynb?ref=datascientyst.com) And the game plan: ![gameplan](https://datascientyst.com/content/images/2022/12/gameplan.webp) ## Step 1: Install Required Modules In this tutorial we need 2 libraries: ```python import pandas as pd from imdb import Cinemagoer ``` The library `Cinemagoer` can be installed by `pip`: ```python pip install cinemagoer ``` Cinemagoer (ex IMDbPY) is a Python package for retrieving the data of the IMDb movie database about movies, people and companies. You can read more for this package here: [cinemagoer.readthedocs](https://cinemagoer.readthedocs.io/en/latest/?ref=datascientyst.com) ## Step 2: Scrape IMDb Movie Data Usually scraping or data collection is a tedious and hard process. Using Python libraries like Cinemagoer can make our life much easier. So with a few lines of code we can extract well structured movie data from IMDb. In this step you can find how to easily search and extract movie data from IMDb. ### Get IMDb movie info Let's start by collecting information about movies by using the IMDb identifier. So for the Game of Thrones - [https://www.imdb.com/title/tt3728462/](https://www.imdb.com/title/tt3728462/?ref=datascientyst.com) we get id - `3728462`. To extract director name we can use following code: ```python from imdb import Cinemagoer # create an instance of the Cinemagoer class ia = Cinemagoer() # get a movie and print its director(s) the_matrix = ia.get_movie('3728462') for director in the_matrix['directors']: print(director['name']) ``` which returns: > Michael Dixon ### Search for movies To search for IMDb movies using python we can use method: `search_movie()` and provide movie title: ```python # search for movie movies = ia.search_movie('Game of Thrones') movies ``` the result is list of movies found from IMDb: ``` [, , , , ,... ``` ### Extract IMDb series In this step we will extract series and episodes for a given movie: ``` [, ``` by using the ID of the movie from the previous step: ```python series = ia.get_movie('0944947') ia.update(series, 'episodes') sorted(series['episodes'].keys()) ``` ### Collect episode data Finally we will collect episodes and their data: - rating - votes - year - plot ```python import pandas as pd from pandas import json_normalize ep_data = [] ls = series.get('episodes') for l in ls.values(): for i in l.values(): # print(i.movieID, '-',i.get('rating'), '-', i, ) data = i.data if 'episode of' in data.keys(): data.pop('episode of') df_temp = pd.DataFrame.from_records([data]) ep_data.append(df_temp) df = pd.concat(ep_data) df ``` The result is DataFrame with data for all episodes from **Game of Thrones**: | | title | kind | season | episode | rating | votes | original air date | | - | ------------------------------------- | ------- | ------ | ------- | -------- | ----- | ----------------- | | 0 | Winter Is Coming | episode | 1 | 1 | 8.901235 | 49519 | 17 Apr. 2011 | | 0 | The Kingsroad | episode | 1 | 2 | 8.601235 | 37465 | 24 Apr. 2011 | | 0 | Lord Snow | episode | 1 | 3 | 8.501235 | 35445 | 1 May 2011 | | 0 | Cripples, Bastards, and Broken Things | episode | 1 | 4 | 8.601235 | 33707 | 8 May 2011 | | 0 | The Wolf and the Lion | episode | 1 | 5 | 9.001235 | 35046 | 15 May 2011 | ## Step 3: Data Processing & Cleaning In this step we would like to transform the original DataFrame data to Season vs Episode data. For this purpose we will use pandas method `pd.pivot_table()`: ```python pd.pivot_table(df, index='episode', columns='season', values='rating').round(1).fillna('').astype(str) ``` the result is table of episodes vs seasons information: | season | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | | ------- | --- | --- | --- | --- | --- | --- | --- | --- | | episode | | | | | | | | | | 1 | 8.9 | 8.6 | 8.6 | 9.0 | 8.3 | 8.4 | 8.5 | 7.6 | | 2 | 8.6 | 8.4 | 8.5 | 9.7 | 8.4 | 9.3 | 8.8 | 7.9 | | 3 | 8.5 | 8.7 | 8.7 | 8.7 | 8.4 | 8.6 | 9.1 | 7.5 | | 4 | 8.6 | 8.6 | 9.5 | 8.7 | 8.5 | 9.0 | 9.7 | 5.5 | | 5 | 9.0 | 8.6 | 8.9 | 8.6 | 8.5 | 9.7 | 8.7 | 6.0 | | 6 | 9.1 | 8.9 | 8.7 | 9.7 | 7.9 | 8.3 | 9.0 | 4.0 | | 7 | 9.1 | 8.8 | 8.6 | 9.0 | 8.8 | 8.5 | 9.4 | | | 8 | 8.9 | 8.6 | 8.9 | 9.7 | 9.8 | 8.3 | | | | 9 | 9.6 | 9.6 | 9.9 | 9.6 | 9.4 | 9.9 | | | | 10 | 9.4 | 9.3 | 9.0 | 9.6 | 9.1 | 9.9 | | | How the code above work: - we select the main parts of the `pivot_table`: - `index='episode'` - `columns='season'` - `values='rating'` - next we round up to 1 decimal point - replace NaN values by empty string If you like to rename the column and index names we can use the following code: ```python df_p = pd.pivot_table(df, index='episode', columns='season', values='rating') df_p = df_p.rename_axis('e') df_p = df_p.rename_axis('s', axis=1) df_p.style.background_gradient(cmap='GnBu', axis=None).format( precision=1, na_rep='') ``` ## Step 4: Visualize IMDb Data in Python To make a beautiful heatmap from the IMDb data we can use Pandas Stylers - `.style.background_gradient()`. ```python pd.pivot_table(df, index='episode', columns='season', values='rating').style.background_gradient(cmap='GnBu', axis=None).format( precision=1, na_rep='') ``` So this will produce heatmap from the series ratings: ![dataisbeautiful-the-absolute-quality-of-game-of-trones](https://datascientyst.com/content/images/2022/10/dataisbeautiful-the-absolute-quality-of-game-of-trones.webp) To make it work we use: - `axis=None` \- in order to apply the heat map over the whole DataFrame - we can apply it to rows - 0 or columns - 1 - `cmap='GnBu'` \- selecting the color styles. To view different options you can provide wrong value and check the error message - `cmap='xxx'` - `precision=1` \- control decimal points for the float numbers - `na_rep=''` \- replace missing values by empty spaces ## Step 5: Create Final Visualization Finally we can do artwork by getting free images from: [pixabay - golden dragon](https://pixabay.com/bg/images/search/golden%20dragon/?ref=datascientyst.com). I'm using Inkscape to create the final image: ![got_viz_watermark](https://datascientyst.com/content/images/2022/12/got_viz_watermark.webp) and one more: ![dataisbeautiful-game-of-thrones](https://datascientyst.com/content/images/2022/10/dataisbeautiful-game-of-thrones.webp) ## Conclusion In this article, we saw how to scrape and transform IMDb data in order to produce popular visualization. We covered different Python libraries and techniques for data collection, wrangling and presenting data. A video with all steps will be published on the Youtube channel. Next we will cover how to create a more popular visualization. If you have ideas for other visualizations from DataIsBeautiful or other places - please suggest them. Happy visualizing! ### 424-dataisbeautiful URL: https://datascientyst.com/424-dataisbeautiful/ Last updated: 2022-10-29T08:35:55.000Z DataIsBeautiful ### How to Extract Domain from URL in Pandas URL: https://datascientyst.com/extract-domain-from-url-in-pandas/ Last updated: 2022-10-28T18:02:03.000Z In this short guide, I'll show you how **to extract domain from a URL column in Pandas DataFrame.** You can also find how to extract netloc, schema, path, params. So at the end you will get: ``` ['https://www.datascientyst.com/cheatsheet','https://www.softhints.com/python'] ``` to: ``` 0 (https, www.datascientyst.com, /cheatsheet, , , ) 1 (https, www.softhints.com, /python, , , ) Name: urls, dtype: object ``` or extracting only domains: ``` 0 www.datascientyst.com 1 www.softhints.com Name: urls, dtype: object ``` ## Setup Let's have DataFrame with URL column from which we will extract list of domains: ```python import pandas as pd data = {'urls': ['https://www.datascientyst.com/cheatsheet','https://www.softhints.com/python']} df = pd.DataFrame(data) ``` DataFrame looks like: | | urls | | - | ---------------------------------------- | | 0 | https://www.datascientyst.com/cheatsheet | | 1 | https://www.softhints.com/python | ## Step 1: Extract domain from URL - urlparse First way to **extract domain from URL in Python** is library - `urlparse`: ```python from urllib.parse import urlparse df['urls'].apply(urlparse) ``` This will extract all information as Series of tuples: ``` 0 (https, www.datascientyst.com, /cheatsheet, , , ) 1 (https, www.softhints.com, /python, , , ) Name: urls, dtype: object ``` To extract only the netloc or the domain we can use: ```python df['urls'].apply(lambda x: urlparse(x)[1]) ``` extracted netloc-s from the URL column: ``` 0 www.datascientyst.com 1 www.softhints.com Name: urls, dtype: object ``` We can extract the full ParseResult from the urlparse library by: ```python df['paths'] = df['urls'].apply(lambda x: urlparse(x)) df['paths'].to_dict() ``` result is full URL information: ``` {0: ParseResult(scheme='https', netloc='www.datascientyst.com', path='/cheatsheet', params='', query='', fragment=''), 1: ParseResult(scheme='https', netloc='www.softhints.com', path='/python', params='', query='', fragment='')} ``` ## Step 2: Extract domain from URL - regex We can use **regular expression in order to extract patterns from the URL columns**. Pandas offers `.str.extract` method: ```python df['urls'].str.extract(r'https://(.*)/') ``` would extract the domains plus the subdomains if any: | | 0 | | - | --------------------- | | 0 | www.datascientyst.com | | 1 | www.softhints.com | Or we can match up to symbol without extracting the symbol. In this case we will search for anything until we match char `t`: ```python df['urls'].str.extract(r'(https.*)(?:t)').head() ``` result: | | 0 | | - | --------------------------------------- | | 0 | https://www.datascientyst.com/cheatshee | | 1 | https://www.softhints.com/py | ## Conclusion We saw two different ways how to **parse URL information with Python and Pandas**. We can extract from Pandas DataFrame information like: - scheme='https' - netloc='www.datascientyst.com' - path='/cheatsheet' - params='' - query='' - fragment='' ![](https://datascientyst.com/content/images/2022/10/extract-domain-from-url-in-pandas.webp) ### Best Data Analysis Libraries for Data Science - Python URL: https://datascientyst.com/best-python-libraries-for-data-analysis-python/ Last updated: 2023-10-31T10:00:20.000Z In this tutorial, we'll discuss the **best libraries for Exploratory Data Analysis in Python**. We will cover these EDA libraries: | Library | GitHub Stars | Contributors | Used by | | ---------------- | ------------ | ------------ | ------- | | pandas-profiling | 9700 | 79 | 9200 | | D-Tale | 3700 | 22 | 501 | | Sweetviz | 2200 | 5 | n/a | | DataPrep | 1400 | 33 | n/a | | AutoViz | 968 | 13 | 265 | | dabl | 684 | 23 | n/a | | klib | 331 | 8 | n/a | ## 1\. Overview Exploratory Data Analysis is a crucial step in the Data Science process. It not only improves quality and consistency of the data, but it also reveals hidden trends and insights. In this article we will use the following DataFrames: ```python import pandas as pd file = 'https://raw.githubusercontent.com/softhints/Pandas-Exercises-Projects/main/data/food_recipes.csv' df = pd.read_csv(file, low_memory=False) file_m = 'https://raw.githubusercontent.com/softhints/Pandas-Exercises-Projects/main/data/movies_metadata.csv' df_m = pd.read_csv(file_m, low_memory=False) ``` You can learn more by: - opening the notebook from: - [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/softhints/Pandas-Exercises-Projects/blob/main/project/Exploratory%20Data%20Analysis%20-%20Python%20Libraries.ipynb?ref=datascientyst.com) - GitHub - [Exploratory Data Analysis - Python Libraries.ipynb](https://github.com/softhints/Pandas-Exercises-Projects/blob/main/project/Exploratory%20Data%20Analysis%20-%20Python%20Libraries.ipynb?ref=datascientyst.com) - watching video: ## 2\. sweetviz - quick summary SweetViz generates beautiful and detailed reports with visualizations. The report is available as HTML output. ### Resources [sweetviz](https://pypi.org/project/sweetviz/?ref=datascientyst.com) > In-depth EDA (target analysis, comparison, feature analysis, correlation) in two lines of code! - `pip install sweetviz` [github - sweetviz](https://github.com/fbdesignpro/sweetviz?ref=datascientyst.com) ### Features - Target analysis - Visualize and compare - Mixed-type associations - Type inference - Summary information ### Code ```python import sweetviz as sv my_report = sv.analyze(df) my_report.show_html() ``` ![python-exploratory-data-analysis-sweetviz](https://datascientyst.com/content/images/2022/10/python-exploratory-data-analysis-sweetviz.gif) We can get quick summary of data to find: - trends - missing values - correlations - categorical data We can see a summary of the data types. Then we can see the associations. There are 3 types of info depending on the column type: - text - numeric - categorical ## 3\. autoviz - visualization As the name suggests it will automatically visualize datasets. One difference to the first one is that: - it will take file as an input - so you need to provide file path and separator - It's a bit slower than previous one Reports are generated by using Bokeh as Jupyter output. ### Resources [autoviz](https://pypi.org/project/autoviz/?ref=datascientyst.com) > Automatically Visualize any dataset, any size with a single line of code. Now you can save these interactive charts as HTML files automatically with the "html" setting. - `pip install autoviz` [github - autoviz](https://github.com/AutoViML/AutoViz?ref=datascientyst.com) ### Features - Visualize and compare - Summary information - Outliers and missing values - Data cleaning improvement suggestions ### Code ```python from autoviz.AutoViz_Class import AutoViz_Class AV = AutoViz_Class() #EDA using Autoviz dft = AV.AutoViz( file, sep=",") ``` ![python-exploratory-data-analysis-autoviz](https://datascientyst.com/content/images/2022/10/python-exploratory-data-analysis-autoviz.gif) First we can see the summary. Number of unique values, missing ones and dtypes. We know that: **prep\_time and cook\_time** are recognized as string columns. But they can be converted to numerical - by replacing ' M'. This library offers - Data cleaning improvement suggestions. This is very useful for beginners. We can see different plots depending on the dtype like: - distribution plot - box plot - heatmaps - bar plots for continuous data - and finally word clouds for text data. ## 4\. pandas-profiling - reports pandas-profiling is the most popular and most used library according to Github. It has a big list of features and it's very fast. There are 7 tabs in the final report. ### Resources [pandas-profiling](https://pypi.org/project/pandas-profiling/?ref=datascientyst.com) > pandas-profiling generates profile reports from a pandas DataFrame. pandas-profiling extends pandas DataFrame with df.profile\_report(), which automatically generates a standardized univariate and multivariate report for data understanding. - `pip install pandas-profiling` [github - profiling](https://github.com/ydataai/pandas-profiling?ref=datascientyst.com) ### Features - **Type inference**: detect the types of columns in a DataFrame - **Essentials**: type, unique values, indication of missing values - **Quantile statistics**: minimum value, Q1, median, Q3, maximum, range, interquartile range - **Descriptive statistics**: mean, mode, standard deviation, sum, median absolute deviation, coefficient of variation, kurtosis, skewness - **Most frequent and extreme values** - **Histograms**: categorical and numerical - **Correlations**: high correlation warnings, based on different correlation metrics (Spearman, Pearson, Kendall, Cramér’s V, Phik) - **Missing values**: counts, matrix, heatmap and dendrograms - **Duplicate rows**: list of the most common duplicated rows - **Text analysis**: most common categories (uppercase, lowercase, separator), scripts (Latin, Cyrillic) and blocks (ASCII, Cyrilic) - **File and Image analysis**: ### Code ```python from pandas_profiling import ProfileReport profile = ProfileReport(df, explorative=True) profile ``` ![python-exploratory-data-analysis-pandas-profiling](https://datascientyst.com/content/images/2022/10/python-exploratory-data-analysis-pandas-profiling.gif) #### Overview In first Tab we can see - overview of data. - alerts - info for the analysis #### Variables - depending on the column type we see different information - we can get more details by toggle - URL analysis - pretty useful feature - netlocs - schema - text data - min and max length - most common words - lower case data - line breaks etc - language detection - most frequent char per script - numeric - stats like: 0-s, negative values, mean, max etc - histogram - common values extreme values / possible outliers #### Interactions We can find relation between different numeric columns #### Correlations There are several methods to get Correlations. We can get a description for each method by clicking on toggle. #### Missing values Missing values can be found in several different ways: - count of null values - matrix - heatmap - dendrogram There is information below each graph. #### Sample In final tab we can see samples from the start and the end of this dataframe ## 5\. dataprep - simple EDA dataprep is advertised as: > The easiest way to prepare data in Python. Some functionalities of DataPrep are inspired by Pandas Profiling. It generates a beautiful profile report from a DataFrame with the create\_report function. Works with Pandas and Dask. ### Resources [dataprep](https://pypi.org/project/dataprep/?ref=datascientyst.com) > DataPrep lets you prepare your data using a single library with a few lines of code. - `pip install -U dataprep` [github - dataprep](https://github.com/sfu-db/dataprep?ref=datascientyst.com) ### Features - Collect data from common data sources (through dataprep.connector) - Do your exploratory data analysis (through dataprep.eda) - Clean and standardize data (through dataprep.clean) ### Code ```python from dataprep.datasets import load_dataset from dataprep.eda import create_report # df = load_dataset("titanic") create_report(df).show_browser() ``` ![python-exploratory-data-analysis-dataprep](https://datascientyst.com/content/images/2022/10/python-exploratory-data-analysis-dataprep.gif) Advantages of this library are: - it's very fast due to highly optimized Dask-based computing module - supports big data - 140+ functions designed for cleaning and validating data - clean\_country - validate\_country - plot\_correlation - plot\_missing To list dataprep functions we can use next code snippets: ```python from dataprep.eda import __all__ print(__all__) ``` which results into: ``` ['plot_correlation', 'compute_correlation', 'render_correlation', 'compute_missing', 'render_missing', 'plot_missing', 'plot', 'compute', 'render', 'DType', 'Categorical', 'Nominal', 'Ordinal', 'Numerical', 'Continuous', 'Discrete', 'DateTime', 'Text', 'create_report', 'create_db_report', 'create_diff_report', 'plot_diff', 'compute_diff', 'render_diff'] ``` and ```python from dataprep.clean import __all__ print(__all__) ``` To get: ``` ['clean_lat_long', 'validate_lat_long', 'clean_email', 'validate_email', 'clean_country', 'validate_country', 'clean_url', 'validate_url', 'clean_phone', 'validate_phone', 'clean_json', 'validate_json', 'clean_ip', 'validate_ip', 'clean_headers', 'clean_address', 'validate_address', 'clean_date', 'validate_date', ``` ## 6\. dabl - single column Dabl focuses less on statistical measures of individual columns, and more on providing a quick overview via visualizations, as well as convenient preprocessing. It's actively developed and not recommended for production. The goal of dabl is to provide handy tool for beginners which build machine learning modules. ### Resources [dabl](https://pypi.org/project/dabl/?ref=datascientyst.com) > Data Analysis Baseline Library. - `pip install dabl` [dabl - github](https://github.com/dabl/dabl?ref=datascientyst.com) ### Features - Analyze single columns - Grouped univariate histograms - Scatter plot for categories - Determine a good grid shape for subplots - Create a mosaic plot from a dataframe - Plots for categorical features in classification - Visualize coefficients of a linear model ### Code ```python import dabl dabl.plot(df, target_col="rating") ``` ![python-exploratory-data-analysis-dabl](https://datascientyst.com/content/images/2022/10/python-exploratory-data-analysis-dabl.png) ## 7\. dtale - interactive D-Tale combines Flask back-end and a React front-end to bring an easy way to view & analyze Pandas data structures. It's interactive and works with JupyterNotebook and JupyterLab. Currently this tool supports such Pandas objects as DataFrame, Series, MultiIndex. [dtale](https://pypi.org/project/dtale/?ref=datascientyst.com) > Data Analysis Baseline Library. - `pip install dtale` [dtale - github](https://github.com/man-group/dtale?ref=datascientyst.com) ### Features - Summarize Data - Duplicates detection - Missing Analysis - Outlier Detection - Custom Filter - Network Viewer - Correlations - Predictive Power Score - Heat Map - Load Data & Sample Datasets ### Code ```python import dtale import pandas as pd dtale.show(df) ``` ![python-exploratory-data-analysis-dtale](https://datascientyst.com/content/images/2022/10/python-exploratory-data-analysis-dtale.gif) ## 8\. klib - rich features klib is a library for importing, cleaning, analyzing and preprocessing data. Functions are divided in two areas: - describe - clean [klib](https://pypi.org/project/klib/?ref=datascientyst.com) > Customized data preprocessing functions for frequent tasks.. - `pip install klib` [klib - github](https://github.com/akanz1/klib?ref=datascientyst.com) ### Features **klib.describe** \- functions for visualizing datasets - `klib.cat_plot(df)` \- returns a visualization of the number and frequency of categorical features - `klib.corr_mat(df)` \- returns a color-encoded correlation matrix - `klib.corr_plot(df)` \- returns a color-encoded heatmap, ideal for correlations - `klib.dist_plot(df)` \- returns a distribution plot for every numeric feature - `klib.missingval_plot(df)` \- returns a figure containing information about missing values **klib.clean** \- functions for cleaning datasets - `klib.data_cleaning(df)` \- performs data cleaning (drop duplicates & empty rows/cols, adjust dtypes,...) - `klib.clean_column_names(df)` \- cleans and standardizes column names, also called inside data\_cleaning() - `klib.convert_datatypes(df)` \- converts existing to more efficient dtypes, also called inside data\_cleaning() - `klib.drop_missing(df)` \- drops missing values, also called in data\_cleaning() - `klib.mv_col_handling(df)` \- drops features with high ratio of missing values based on informational content - `klib.pool_duplicate_subsets(df)` \- pools subset of cols based on duplicates with min. loss of information ### Code ```python import klib klib.missingval_plot(df_m) df_cleaned = klib.data_cleaning(df_m) klib.corr_plot(df_m) klib.corr_plot(df_cleaned, target='revenue') klib.dist_plot(df_m) klib.corr_mat(df_cleaned) ``` ![python-exploratory-data-analysis-klib-corr](https://datascientyst.com/content/images/2022/10/python-exploratory-data-analysis-klib-corr.png) ![python-exploratory-data-analysis-klib-dist](https://datascientyst.com/content/images/2022/10/python-exploratory-data-analysis-klib-dist.png) ## 9\. Conclusion In this article, we explored some of the **best libraries of Exploratory Data Analysis and their features in the Python ecosystem**. Using best libraries can help in many aspects of data science process: - speed up data science projects - make the optimal decisions - clean errors and confusion If you like to learn more about data exploration process please check previous article: [Exploratory Data Analysis Python and Pandas with Examples](https://datascientyst.com/exploratory-data-analysis-pandas-examples/) ### How to Convert Pandas DataFrame to Dictionary URL: https://datascientyst.com/convert-a-pandas-dataframe-to-a-dictionary/ Last updated: 2022-10-19T15:37:28.000Z In this short guide, I'll show you how to **convert Pandas DataFrame to dictionary**. You can also find how to use Pandas method - `to_dict()`. So at the end you will get from **DataFrame to Python dict**: | | day | numeric | | - | --- | ------- | | 0 | 1 | 1 | | 1 | 2 | 2 | | 2 | 3 | 3 | | 3 | 4 | 4 | | 4 | 5 | 5 | to: ``` {'day': {0: 1, 1: 2, 2: 3, 3: 4, 4: 5, 5: 6}, 'numeric': {0: 1, 1: 2, 2: 3, 3: 4, 4: 5, 5: 6}} ``` or any other dict-like format. We will also cover the following examples: - `dict` (default) : dict like {column -> {index -> value}} - `list` : dict like {column -> \[values\]} - `series` : dict like {column -> Series(values)} - `split` : dict like - `{'index' -> [index], 'columns' -> [columns], 'data' -> [values]}` - `tight` : dict like `{'index' -> [index], 'columns' -> [columns], 'data' -> [values], 'index_names' -> [index.names], 'column_names' -> [column.names]}` - `records` : list like - `[{column -> value}, ... , {column -> value}]` - `index` : dict like {index -> {column -> value}} To start, here is the syntax that you may apply in order to convert DataFrame to dict: ```python df.to_dict() ``` In the next section, I'll review the steps to apply the above syntax in practice. ## Step 1: Create a DataFrame Lets create a DataFrame which has a two columns: ```python import pandas as pd data={'day': [1, 2, 3, 4, 5, 6], 'numeric': [1, 2, 3, 4, 5, 6]} df = pd.DataFrame(data) df ``` result: | | day | numeric | | - | --- | ------- | | 0 | 1 | 1 | | 1 | 2 | 2 | | 2 | 3 | 3 | | 3 | 4 | 4 | | 4 | 5 | 5 | ## Step 2: DataFrame to dict - {column -> {index -> value}} In order to extract DataFrame as Python dictionary we need just this line: ```python df.to_dict() ``` result: ``` {'day': {0: 1, 1: 2, 2: 3, 3: 4, 4: 5, 5: 6}, 'numeric': {0: 1, 1: 2, 2: 3, 3: 4, 4: 5, 5: 6}} ``` By default method `to_dict()` use as parameter - `orient='list'` and will produce dict form of: ``` {column -> {index -> value}} ``` ## Step 3: DataFrame to dict - list - {column -> \[values\]} What if you like to get a dictionary only with the values? In this case we will use `orient='list'` in order to exclude index from the output dictionary: ```python df.to_dict(orient='list') ``` ``` {'day': [1, 2, 3, 4, 5, 6], 'numeric': [1, 2, 3, 4, 5, 6]} ``` Note: have in mind that there is a difference between Python dict and JSON format. ## Step 3: split - {'index' -> \[index\], 'columns' -> \[columns\], 'data' -> \[values\]} Another form of dictionary which includes: - index - columns - data as separate lists can be extracted by: ```python df.to_dict(orient='split') ``` ``` {'index': [0, 1, 2, 3, 4, 5], 'columns': ['day', 'numeric'], 'data': [[1, 1], [2, 2], [3, 3], [4, 4], [5, 5], [6, 6]]} ``` and then: ```python df['yyyy'].astype(str) + '-'+ df['mm'].astype(str) ``` ## Step 4: Convert DataFrame to dict - tight format If you like to get a tight format of DataFrame as dict we can use the parameter `orient='tight'`. This option includes the label names - index and column names in the output dict: ```python df.to_dict(orient='tight') ``` to get: {'index': \[0, 1, 2, 3, 4, 5\], 'columns': \['day', 'numeric'\], 'data': \[\[1, 1\], \[2, 2\], \[3, 3\], \[4, 4\], \[5, 5\], \[6, 6\]\], 'index\_names': \[None\], 'column\_names': \[None\]} ## Step 5: Convert DataFrame to dict - records If we like to extract only the values and the column labels as dictionary we can do it by - `orient='records'`: ```python df.to_dict(orient='records') ``` We will get the list of dicts from the original DataFrame: ``` [{'day': 1, 'numeric': 1}, {'day': 2, 'numeric': 2}, {'day': 3, 'numeric': 3}, {'day': 4, 'numeric': 4}, {'day': 5, 'numeric': 5}, {'day': 6, 'numeric': 6}] ``` ## Step 6: Convert DataFrame to dict - index Finally we can convert DataFrame to dictionary with index by: ```python df.to_dict(orient='index') ``` result: ``` {0: {'day': 1, 'numeric': 1}, 1: {'day': 2, 'numeric': 2}, 2: {'day': 3, 'numeric': 3}, 3: {'day': 4, 'numeric': 4}, 4: {'day': 5, 'numeric': 5}, 5: {'day': 6, 'numeric': 6}} ``` More information about method - [to\_dict](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to%5Fdict.html?ref=datascientyst.com) ![](https://datascientyst.com/content/images/2022/10/convert-a-pandas-dataframe-to-a-dictionary.png) ### ValueError: DataFrame constructor not properly called! - Pandas URL: https://datascientyst.com/valueerror-dataframe-constructor-not-properly-called-pandas/ Last updated: 2022-10-18T06:27:17.000Z In this tutorial, we'll take a look at the Pandas error: ``` ValueError: DataFrame constructor not properly called! ``` First, we'll create examples of how to produce it. Next, we'll explain the reason and finally, we'll see how to fix it. ## ValueError: DataFrame constructor not properly called Let's try to create DataFrame by: ```python import pandas as pd df = pd.DataFrame(0) df = pd.DataFrame('a') ``` All of them result into error: > ValueError: DataFrame constructor not properly called! The same will happen if we try to create DataFrame from directly: ```python class k: pass a = k() pd.DataFrame(a) ``` ## Reason There are multiple reasons to get error like: > ValueError: DataFrame constructor not properly called! ### Pass single value One reason is trying to pass single value and no index: ```python df = pd.DataFrame('a') ``` this results in: ``` ValueError: DataFrame constructor not properly called! ``` ### Object to DataFrame Sometimes we need to convert object as follows: ```python class k: def __init__(self, name, ): self.name = name self.num = 0 a = k('test') pd.DataFrame(a) ``` This is not possible and result in: ``` ValueError: DataFrame constructor not properly called! ``` ## Solution - single value To solve error - **ValueError: DataFrame constructor not properly called** \- when we pass a single value we need to use a list - add square brackets around the value. This will convert the input vector value: ```python import pandas as pd df = pd.DataFrame(['a']) ``` The result is successfully created DataFrame: | | 0 | | - | - | | 0 | a | ## Solution - object to DataFrame To convert object to DataFrame we can follow next steps: ```python class k: def __init__(self, name, ): self.name = name self.num = 0 a = k('test') ``` To convert object "a" to DataFrame first convert it to JSON data by: ```python import json print(json.dumps(a.__dict__)) ``` the result is: ``` '{"name": "test", "num": 0}' ``` Next we convert the JSON string to dict: ```python json.loads(data) ``` result: ``` {'name': 'test', 'num': 0} ``` And finally we create DataFrame by: ```python pd.DataFrame.from_dict(json.loads(data), orient='index').T ``` the object is converted to Pandas DataFrame: | | name | num | | - | ---- | --- | | 0 | test | 0 | ## Conclusion In this article, we discussed error: **ValueError: DataFrame constructor not properly called!** reasons and possible solutions. ### Free Public Datasets for Data Science Projects URL: https://datascientyst.com/datasets/ Last updated: 2026-02-20T22:04:03.000Z In this post we can **find free public datasets for Data Science projects.** There is a big number of datasets which cover different areas - machine learning, presentation, data analysis and visualization. You can find information for: - **Data sources** \- big datasets collections which has curated data and advanced searching - Sample datasets - datasets good for data analysis. Starting point for beginners who would like to learn Data Science - **Datasets resources** \- useful resources for datasets which can be loaded easily - **Datasets from Python libraries** \- load datasets with single line of code from different Python libraries like 'seaborn' **Note** This post will be updated on regular basis so please suggest new ideas and datasets in the comment section below. ## 0\. List of datasets for machine-learning research You can find big currated list with different datasets on: [List of datasets for machine-learning research](https://en.wikipedia.org/wiki/List%5Fof%5Fdatasets%5Ffor%5Fmachine-learning%5Fresearch?ref=datascientyst.com): - Image data - Text data - Sound data - Signal data - Chemical data - Physical data - Biological data ## 1\. Dataset Sources Below we can find a table of dataset collections. Most of them have advanced searching by: - file type - size - number of rows - tags | # | site | description | | - | -------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- | | 1 | [https://www.kaggle.com/datasets/](https://www.kaggle.com/datasets/?ref=datascientyst.com) | [https://www.kaggle.com/docs/datasets](https://www.kaggle.com/docs/datasets?ref=datascientyst.com) | | 2 | [https://datasetsearch.research.google.com/](https://datasetsearch.research.google.com/?ref=datascientyst.com) | Google Dataset Search | | 3 | [https://azure.microsoft.com/en-us/services/open-datasets/](https://azure.microsoft.com/en-us/services/open-datasets/?ref=datascientyst.com) | Azure Open Datasets | | 4 | [https://www.openml.org/search?type=data](https://www.openml.org/search?type=data&ref=datascientyst.com) | 4325 datasets found (verified) | | 5 | [https://grouplens.org/datasets/](https://grouplens.org/datasets/?ref=datascientyst.com) | several datasets | | 6 | [https://datahub.io/search](https://datahub.io/search?ref=datascientyst.com) | thousands of datasets | Other soruces: - [World Bank Open Data](https://data.worldbank.org/?ref=datascientyst.com) ## 2\. Datasets samples There are several listed below which are used in this site for demonstration of data science basics: | # | dataset | size | link | description | | | | | | | - | ------------------ | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------- | | | | | | | 1 | the-movies-dataset | 45466, 24 | [https://www.kaggle.com/datasets/rounakbanik/the-movies-dataset](https://www.kaggle.com/datasets/rounakbanik/the-movies-dataset?ref=datascientyst.com) | [https://grouplens.org/datasets/movielens/latest/](https://grouplens.org/datasets/movielens/latest/?ref=datascientyst.com) | | | | | | | 2 | Food Recipes | 8009, 16 | [https://www.kaggle.com/datasets/sarthak71/food-recipes](https://www.kaggle.com/datasets/sarthak71/food-recipes?ref=datascientyst.com) | | | | | | | | 3 | | | | | | | | | | | 4 | | | | | | | | | | | 5 | | | | | | | | | | | 6 | | | | | | | | | | | 7 | | | | | | | | | | ## 3\. Datasets resources - [93 Datasets That Load With A Single Line of Code](https://towardsdatascience.com/93-datasets-that-load-with-a-single-line-of-code-7b5ffe62b655?ref=datascientyst.com) - [Wikidata - the free knowledge base with 99,946,076 data items](https://www.wikidata.org/wiki/Wikidata:Main%5FPage?ref=datascientyst.com) - [UN Data](https://data.un.org/datamartinfo.aspx?ref=datascientyst.com) ## 4\. Read Kaggle Datasets To read Kaggle datasets we can use the Python library `kaggle`. Downloading dataset from kaggle with Python code is available from method: `dataset_download_file`: ```python import kaggle kaggle.api.authenticate() kaggle.api.dataset_download_file('dorianlazar/medium-articles-dataset', file_name='medium_data.csv', path='data/') ``` For more information and examples refer to: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) ## 5\. Load Datasets by Python libraries In this section we can find several useful datasets for different purposes like: - machine learning - visualization - testing - creating own datasets with fake data ### 5.1 datasets - machine learning Python library `datasets` offers a huge number of free and easy to use datasets. It can be installed by: ```bash pip install datasets ``` To list all available datasets we can use method: `datasets.list_datasets()`: ```python from datasets import list_datasets, load_dataset print(list_datasets()) ``` It will return more than 7000 datasets. To load dataset we can use method: `datasets.load_dataset(dataset_name, **kwargs)`: ```python squad_dataset = load_dataset('squad') squad_dataset ``` This give us two datasets: - training dataset - validation dataset DatasetDict({ train: Dataset({ features: \['id', 'title', 'context', 'question', 'answers'\], num\_rows: 87599 }) validation: Dataset({ features: \['id', 'title', 'context', 'question', 'answers'\], num\_rows: 10570 }) }) To access the dataset for training we can use: `squad_dataset['train']`. Finally we can loaded as Pandas DataFrame by: ```python import pandas as pd pd.DataFrame(squad_dataset['train']) ``` ### 5.2 pandas - test datasets Pandas offers multiple ways to download datasets with a single line of code. Let's cover few of them starting with the test data in Pandas github: - [Pandas data files - csv, xml, html](https://github.com/pandas-dev/pandas/tree/main/pandas/tests/io/data?ref=datascientyst.com) - [tips.csv](https://raw.githubusercontent.com/pandas-dev/pandas/master/pandas/tests/io/data/csv/tips.csv?ref=datascientyst.com) #### scrape wiki tables Next we can load data from Pandas by scraping wikipedia: ```python pd.read_html('https://en.wikipedia.org/wiki/Population_growth')[2] ``` #### create datasets We can create random or fake datasets with Pandas by: - [How To Make a Fake Data Set in Python and Pandas](https://datascientyst.com/make-fake-data-set-python-pandas/) - [How to Easily Create Dummy DataFrame with Test Data?](https://datascientyst.com/create-easily-dummy-dataframe-test-data/) - [How to Create a Pandas DataFrame of Random Integers](https://datascientyst.com/how-to-create-a-dataframe-of-random-integers-with-pandas/) ### 5.3 seaborn - visualization datasets Seaborn offers free tests which are good for visualization. With single line of code we can get DataFrame good for data wrangling and visualization: ```python import seaborn as sns df = sns.load_dataset('flights') ``` All datasets available from seaborn library: [seaborn-data](https://github.com/mwaskom/seaborn-data?ref=datascientyst.com). #### sklearn-learn - machine learning We can get sample datasets from `sklearn-learn` by methods like: `load_iris` ```python from sklearn.datasets import load_iris iris = load_iris() ``` To find more sample datasets from sklearn we can use the next code: ```python from sklearn import datasets dir(datasets) ``` This will list all available options like: ``` 'load_sample_images', 'load_svmlight_file', 'load_svmlight_files', 'load_wine', 'make_biclusters', 'make_blobs', 'make_checkerboard', 'make_circles', ``` You can find more about sklearn-learn datasets on this link: [sklearn.datasets: Datasets](https://scikit-learn.org/stable/modules/classes.html?ref=datascientyst.com#module-sklearn.datasets). ### 5.4 dataprep - data analysis To load dataset we can use method: `load_dataset` ```python from dataprep.datasets import load_dataset df = load_dataset("titanic") ``` to list datasets we can use: ```python from dataprep.datasets import get_dataset_names get_dataset_names() ``` which results into several datasets like: ``` ['waste_hauler', 'wine-quality-red', 'countries', 'house_prices_train', 'iris', 'adult', 'covid19', 'titanic', 'patient_info', 'house_prices_test'] ``` More information about dataprep datasets: [Datasets DataPrep](https://docs.dataprep.ai/user%5Fguide/datasets/introduction.html?ref=datascientyst.com) ## 6\. Top 10 sites with interesting datasets The table below is based on this Kaggle list: [Top 10 sites with interesting datasets](https://www.kaggle.com/discussions/general/422290?ref=datascientyst.com) Here's a concise Markdown table: | **Source** | **Best For** | | ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | | [UCI ML Repository](http://archive.ics.uci.edu/datasets?ref=datascientyst.com) | Classic ML datasets for research, education & benchmarking | | [Google Dataset Search](https://datasetsearch.research.google.com/?ref=datascientyst.com) | Discovering free datasets across any domain or niche | | [Data.gov](https://www.data.gov/?ref=datascientyst.com) | U.S. government datasets — health, finance, climate & agriculture | | [World Bank Open Data](https://data.worldbank.org/?ref=datascientyst.com) | Global socio-economic indicators — poverty, education & infrastructure | | [Reddit – r/datasets](https://www.reddit.com/r/datasets/?ref=datascientyst.com) | Community-shared datasets across diverse & unusual topics | | [AWS Public Datasets](https://registry.opendata.aws/?ref=datascientyst.com) | Large-scale datasets — genomics, astronomy & climate data | | [FiveThirtyEight](https://data.fivethirtyeight.com/?ref=datascientyst.com) | Journalism datasets — politics, sports & public opinion | | [Data.gov.uk](https://www.data.gov.uk/?ref=datascientyst.com) | UK government datasets — crime, health & transportation | | [DataHub.io](https://datahub.io/?ref=datascientyst.com) | Open datasets across finance, climate & social sciences | | [Data.world](https://data.world/?ref=datascientyst.com) | Collaborative platform for discovering & sharing datasets | ## 7\. Free Datasets for Data Science Projects Below you can find a table with free **datasets for data science projects** — sorted by approximate dataset count (popularity/size of collection): | **Source** | **What You Get** | **Best For** | | ----------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- | --------------------------------------------------------- | | [Kaggle Datasets](https://www.kaggle.com/datasets?ref=datascientyst.com) | 1,000s of free datasets across finance, health, NLP, images & more | ML models, EDA dashboards, recommendation systems | | [Google Dataset Search](https://datasetsearch.research.google.com/?ref=datascientyst.com) | Indexed millions of datasets from government, academic & community sources | Macro trends, geospatial & economic forecasting | | [OpenML](https://www.openml.org/?ref=datascientyst.com) | 20,000+ ML datasets with rich metadata for benchmarking | Algorithm benchmarking, classification & regression tasks | | [UCI ML Repository](https://archive.ics.uci.edu/ml/index.php?ref=datascientyst.com) | Classic curated datasets (Iris, Adult Income, Wine, etc.) | Classification, clustering & feature engineering practice | | [World Bank Open Data](https://data.worldbank.org/?ref=datascientyst.com) | Global GDP, poverty, education & demographic indicators | Time-series analysis & country comparisons | | [FiveThirtyEight Data](https://data.fivethirtyeight.com/?ref=datascientyst.com) | Journalism-backed datasets on sports, politics & culture | Data storytelling & visual dashboards | | [Mozilla Common Voice](https://commonvoice.mozilla.org/?ref=datascientyst.com) | Free multilingual crowdsourced speech corpus | Speech recognition & NLP models | | [LabelMe](http://labelme.csail.mit.edu/?ref=datascientyst.com) | Annotated image dataset from MIT CSAIL | Computer vision, object detection & segmentation | | [Figshare](https://figshare.com/?ref=datascientyst.com) | Open-access research datasets across scientific fields | Scientific analysis & research visualization | | [DeepDataLake](https://deepdatalake.com/?ref=datascientyst.com) | 25+ free datasets in CSV, JSON & SQL formats | Sales forecasting, customer segmentation & text analytics | ## 8\. Conclusion In this article, we covered **free datasets sources and discussed common ways to download dataset from them**. Through practical examples, we learned how to download and use those datasets in Python and Pandas. We covered different Python libraries which offer public datasets for learning. Finally, we covered how to create test datasets with fake data. Those **datasets and ideas should be sufficient for practicing and learning data science.** ### How to Convert String, DateTime Or TimeStamp to Time in Pandas URL: https://datascientyst.com/convert-string-datetime-timestamp-time-pandas/ Last updated: 2022-10-15T08:04:09.000Z In this article we will see how to **extract time only from string or datetime in Pandas.** First, we'll create an example DataFrame to test it. Next, we'll explain several examples in more detail. **(1) extract time with .dt.time - datetime.time** ```python df['date'].dt.time ``` **(2) get time by .dt.strftime('%H:%M') as string** ```python df['date'].dt.strftime('%H:%M') ``` ## Setup Let's work with the following DataFrame which has date and time information stored as a string: ```python import pandas as pd dict = {'date': {0: '28-01-2022 5:25:00 PM', 1: '27-02-2022 6:25:00 PM', 2: '30-03-2022 7:25:00 PM', 3: '29-04-2022 8:25:00 PM', 4: '31-05-2022 9:25:00 PM'}, 'date_short': {0: 'Jan-2022', 1: 'Feb-2022', 2: 'Mar-2022', 3: 'Apr-2022', 4: 'May-2022'}} df = pd.DataFrame(dict) ``` DataFrame looks like: | | date | date\_short | | - | --------------------- | ----------- | | 0 | 28-01-2022 5:25:00 PM | Jan-2022 | | 1 | 27-02-2022 6:25:00 PM | Feb-2022 | | 2 | 30-03-2022 7:25:00 PM | Mar-2022 | | 3 | 29-04-2022 8:25:00 PM | Apr-2022 | | 4 | 31-05-2022 9:25:00 PM | May-2022 | ## String to datetime Convert string to datetime with Pandas: ```python df['date'] = pd.to_datetime(df['date']) ``` ## Extract time with .dt.time as datetime.time Once we have datetime in Pandas we can extract time very easily by using: `.dt.time`. Default format of extraction is: `HH:MM:SS`. It returns a numpy array of datetime.time objects. ```python df['date'].dt.time ``` will give us: ``` 0 17:25:00 1 18:25:00 2 19:25:00 3 20:25:00 4 21:25:00 Name: date, dtype: object ``` ## Get time by .dt.strftime('%H:%M') as string For custom format we can use - `.dt.strftime('%H:%M')`. `strftime` returns strings formatted in time format given as parameter ```python df['date'].dt.strftime('%H:%M') ``` will give us: ``` 0 17:25 1 18:25 2 19:25 3 20:25 4 21:25 Name: date, dtype: object ``` We can find more information on this link: [pandas.Series.dt.strftime](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.strftime.html?ref=datascientyst.com) ![](https://datascientyst.com/content/images/2022/10/convert-string-datetime-timestamp-time-pandas.png) ## Conclusion To summarize we saw how to extract time with Pandas in two ways. First by using `.dt.time` as datetime. Second one extracts time as a string in any format needed. ### KeyError:0 - Create DataFrame in Pandas URL: https://datascientyst.com/keyerror-0-create-dataframe-pandas/ Last updated: 2022-10-16T05:32:47.000Z In this tutorial, we'll take a look at the Pandas error: ``` KeyError:0 ``` First, we'll create an example of how to reproduce it. Next, we'll explain the reason and finally, we'll see how to fix it. ## Example Let's work with the following DataFrame: ```python import pandas as pd data={'day': [1, 2, 3, 4, 5], 'numeric': [22, 222, '22K', '2M', '0.01 B']} df = pd.DataFrame(data) ``` Data is: | day | numeric | | --- | ------- | | 1 | 22 | | 2 | 222 | | 3 | 22K | | 4 | 2M | | 5 | 0.01 B | ## Reason There are multiple reasons to get error like: > KeyError:0 ### Not existing column One reason is trying to access column which don't exist: ```python df[0] ``` result in: ``` KeyError:0 ``` ### Bad Input - Creating DataFrame with dict Sometimes when we work with API we get dict as input. Some dict elements might cause similar problems. In that case we need to drop elements from the input by: ```python data.pop('episode of') ``` In this case the element 'episode of' is pointing to an instance. Working with IMDB library [cinemagoer](https://pypi.org/project/cinemagoer/?ref=datascientyst.com) return result as: ```python {'title': 'The Iron Throne', 'kind': 'episode', 'episode of': , 'season': 8, 'episode': 6, 'rating': 4.001234567891, 'votes': 249455, 'original air date': '19 May 2019', 'year': '2019', 'plot': "\nIn the aftermath of the devastating attack on King's Landing, Daenerys must face the survivors. "} ``` Notice item: ``` 'episode of': , ``` This item cause error: ``` KeyError:0 ``` ## Solution - wrong column For the wrong column error we can check what are the current columns of the DataFrame: ```python df.columns ``` result: ``` Index(['day', 'numeric'], dtype='object') ``` Then we can access the correct one: ```python df['day'] ``` ## Solution - bad input To investigate and solve bad inputs which cause: ``` KeyError:0 ``` We can try to isolate problematic values. For example creating DataFrame with line: ``` 'episode of': , ``` will cause the error. Dropping this line from the input by: ```python if 'episode of' in data.keys(): data.pop('episode of') ``` will solve the error. ### Pandas Cheat Sheet: Data Cleaning URL: https://datascientyst.com/pandas-cheat-sheet-data-cleaning/ Last updated: 2023-03-17T23:00:55.000Z A practical **Pandas Cheat Sheet: Data Cleaning** useful for everyday working with data. This Pandas cheat sheet contains ready-to-use codes and steps for data cleaning. The **cheat sheet aggregate the most common operations used in Pandas for:** analyzing, fixing, removing - incorrect, duplicate or wrong data. This cheat sheet will act as a **guide for data science beginners** and help them with various fundamentals of data cleaning. Experienced users can use it as a quick reference. - [Data Cleaning Steps with Python and Pandas](https://datascientyst.com/data-cleaning-steps-python-example/) - [Data Cleaning - Kaggle course](https://www.kaggle.com/learn/data-cleaning?ref=datascientyst.com) - [Working with missing data - Pandas Docs](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/missing%5Fdata.html?ref=datascientyst.com) - [Data Cleaning Steps - kaggle](https://www.kaggle.com/getting-started/250322?ref=datascientyst.com) ## EDA Exploratory Data Analysis ![](https://datascientyst.com/content/images/2022/10/eda.png) `df.info()` DataFrame columns, dtypes and memory `df.describe()` Returns columns coverage and types `df.head(7)` Returns first N rows `df.sample(2)` return random samples `df.shape` return DataFrame dimensions `df.columns` returns DataFrame columns ## Duplicates Detect and Remove duplicates ![](https://datascientyst.com/content/images/2022/10/duplicates.png) `df.diet.nunique()` number of unique values in column `df.diet.unique()` unique values in column `df['col_1'].value_counts(dropna=False)` return series of unique values and counts in column `df.duplicated(keep='last')` find duplicates and keep only the last record `df.drop_duplicates(subset=['col_1'])` drop duplicates from column(s) `df[df.duplicated(keep=False)].index` get indexes of all detected duplications: ## Missing values Working with missing data ![](https://datascientyst.com/content/images/2022/10/missing.png) `df.isna()` return True or False for missing values `df['col_1'].notna()` return True or False for non-NA data `df.isna().all() s[s == True]` Columns which contains only NaN values `df.isna().any()` Detect columns with NaN values `df['col_1'].fillna(0)` Fill NaN with string or 0 `import seaborn as sns sns.heatmap(df.isna(),cmap = 'Greens')` plot missing values `s.loc[0] = None s.loc[0] = np.nan` Insert missing data `df.dropna(axis=0)` droping rows with missing data `df.dropna(axis=1, how='any')` Drop columns with NaN values ## Outliers Detect and remove outliers ![](https://datascientyst.com/content/images/2022/10/outliers-1.png) `df['col_1'].describe()` detecting outliers with describe() `import seaborn as sns sns.boxplot(data=df[['col_1', 'col_2']])` detect outliers with boxplot `q_low = df['col'].quantile(0.01) q_hi = df['col'].quantile(0.99) df[(df['col'] < q_hi)&(df['col'] > q_low)]` remove outliers with quantiles `import numpy as np ab = np.abs(df['col']-df['col'].mean()) std = (3*df['col'].std()) df[ab <= std ]` remove outliers with standart deviation ## Wrong data Detect wrong data ![](https://datascientyst.com/content/images/2022/10/wrongdata.png) `df[df['col_1'].str.contains(r'[@#&$%+-/*]')]` Detect special symbols `df[df['col_1'].map(lambda x: x.isascii())]` Detect (non) ascii characters `df['col']\ .loc[~df['col'].str.match(r'[0-9.]+')]` find pattern with regex `import numpy as np np.where(df['col']=='',df['col2'],df['col'])` detect empty spaces `df[df['col_1'].str.contains('[A-Za-z]')]` Detect latin symbols `df.applymap(np.isreal)` detect non numeric rows `df['city'].str.len().value_counts()` Count values by lenght ## Wrong format Detect wrong format ![](https://datascientyst.com/content/images/2022/10/wrongformat.png) `df.apply(pd.to_numeric,errors='coerce')\ .isna().any()` detect wrong numeric format `pd.to_datetime(df['date_col'],errors='coerce')` detect wrong datetime format `import pandas_dedupe dd_df = pandas_dedupe.dedupe_dataframe( df, field_properties=['col1', 'col2'], canonicalize=['col1'], sample_size=0.8 )` Find typos and misspelling - Deduplication and canonicalization with pandas\_dedupe `from difflib import get_close_matches w = ['apes', 'apple', 'peach', 'puppy'] get_close_matches('ape', w, n=3, cutoff=0.8)` Use difflib to find close matches ## Fix errors Fix errors in Pandas ![](https://datascientyst.com/content/images/2022/10/fixerrors.png) `df.convert_dtypes()` Convert the DataFrame to use best possible dtypes `df.astype({'col_1': 'int32'})` Cast col\_1 to int32 using a dictionary `df.fillna(method='ffill')` Propagate non-null values forward or backward `values = {'A': 0, 'B': 1, 'C': 2, 'D': 3} df.fillna(value=values)` Replace all NaN elements in column with dict `df.ffill(axis = 0)` fill the missing values row wise ## Replacing Replace data in DataFrame ![](https://datascientyst.com/content/images/2022/10/replace.png) `df['col'] = df['col'].str.replace(' M', '')` replace string from column `df['col'].str.replace(' M', '')\ .fillna(0).astype(int)` replace and convert column to integer `df['col'].str.replace('A7', '7', regex=False)` Replace values in column - no regex `df.replace(r'\r+|\n+|\t+','', regex=True)` Find and replace line breaks - new line, tab - regex `df['col'].str.replace('\s+', '', regex=True)` Replace multiple white spaces `df['col'].str.rstrip('\r\n')` Replace line breaks from the right `p = r'<[^<>]*>' df['col1'].str.replace(p, '', regex=True)` Replace HTML tags ## Drop Drop rows, columns, index, condition ![](https://datascientyst.com/content/images/2022/10/drop.png) `df.drop('col_1', axis=1, inplace=True)` Drop one column by name `df.drop(['col1', 'col2'], axis=1)` Drop multiple columns by name `df.dropna(axis=1, how='any')` Drop columns with NaN values `df.drop(0)` Drop rows by index - 0 `df.drop([0, 2, 4])` drop multiple rows `df[(df['col1'] > 0) & (df['col2'] != 'open')]` drop rows by condition `df.reset_index()` drop index ## Pandas cheat sheet: data cleaning ![](https://datascientyst.com/content/images/2022/10/cleaning_cheatsheet.png) ## Data cleaning steps Below you can find the data cleaning steps in order to ensure that your dataset is good for decisions: - Are there obvious errors - mixed data - data without structured - different column lengths - Is data consistent - same units - km, m, etc - same way - Paris, Par - same data types - '0.1', 0,1 - Missing values - how many - reasons for missing values - Same format - DD/MM/YYYY, YY-MM-DD - '15 M', '15:00' - Duplicates - duplicates in rows and values - drop duplicates - dedupe - keep only one - Par / Paris / Paris (Fra) - Outliers - detect and remove outliers - Data bias - is data biased - does it represent the whole population or part of it - how data was selected - Data noise - unwanted data items and values - invalid data - Data leakage - use information from outside of the training model Note: Please add ideas and suggestions in the comments below. Thanks you 💕 ![](https://datascientyst.com/content/images/2022/10/outliers-1.png) ### ValueError: All arrays must be of the same length - Pandas URL: https://datascientyst.com/valueerror-all-arrays-must-be-of-the-same-length-pandas/ Last updated: 2022-10-13T05:46:03.000Z In this tutorial, we'll see **how to solve Pandas error**: ``` ValueError: All arrays must be of the same length ``` First, we'll create an example of how to produce it. Next, we'll explain the reason and finally, we'll see how to fix it. ## Example Let's try to create the following DataFrame: ```python import pandas as pd data={'day': [1, 2, 3, 4, 5], 'numeric': [1, 2, 3, 4, 5, 6]} df = pd.DataFrame(data) ``` this would cause: ``` ValueError: All arrays must be of the same length ``` ## Reason The problem is that we try to create DataFrame from arrays with different length: ```python for i in data.values(): print(len(i)) ``` result: ``` 5 6 ``` So we will get this error when we try to create a DataFrame with columns of different lengths. So we need to use equal sized input arrays. ### API errors We can get this error often when we work with different API-s and try to create DataFrames from the results. Investigate the input data and correct it if needed. ## Solution In order to solve the error we can change the array length. A generic solution would be something like: ```python df = pd.DataFrame.from_dict(data, orient='index') df = df.transpose() ``` which will result into DataFrame like: | | day | numeric | | - | --- | ------- | | 0 | 1.0 | 1.0 | | 1 | 2.0 | 2.0 | | 2 | 3.0 | 3.0 | | 3 | 4.0 | 4.0 | | 4 | 5.0 | 5.0 | | 5 | NaN | 6.0 | ### How does it work First we create DataFrame from the existing data as index: ```python pd.DataFrame.from_dict(data, orient='index') ``` this result into: | | 0 | 1 | 2 | 3 | 4 | 5 | | ------- | - | - | - | - | - | --- | | day | 1 | 2 | 3 | 4 | 5 | NaN | | numeric | 1 | 2 | 3 | 4 | 5 | 6.0 | Then we just transpose the results. By using `from_dict(orient='index')` we can have different sized arrays as DataFrame input. ## Note Since error: "ValueError: All arrays must be of the same length" suggest data inconsistency be sure that input data is correct. Check if data is aligned correctly and can be used. ## Conclusion In this article, we saw how to investigate and solve error: "ValueError: All arrays must be of the same length". ### 416-pandas-error URL: https://datascientyst.com/416-pandas-error/ Last updated: 2026-03-02T22:12:03.000Z Pandas Error ### ValueError: If using all scalar values, you must pass an index - Pandas URL: https://datascientyst.com/valueerror-if-using-all-scalar-values-you-must-pass-an-index-pandas/ Last updated: 2023-01-15T08:43:10.000Z In this tutorial, we'll take a closer look at the Pandas error: **"ValueError: If using all scalar values, you must pass an index"** You can find explanation and solution on the image below: ![fix-valueerror-if-using-all-scalar-values-you-must-pass-an-index](https://datascientyst.com/content/images/2023/01/fix-valueerror-if-using-all-scalar-values-you-must-pass-an-index.webp) Quick fixes: **(1) add index** ```python pd.DataFrame(dct, index=[0]) ``` **(2) use vector values** ```python dct = {k:[v] for k,v in dct.items()} ``` **(3) Wrap with list** ```python dct = {'col1': ['val1'], 'col2': [2]} pd.DataFrame(dct, index=[0]) ``` First, we'll create an example of how to produce it. Next, we'll explain the leading cause of the error. And finally, we'll see how to fix it. ## Example Now, let's see an example that generates a Pandas error: > ValueError: If using all scalar values, you must pass an index ```python import pandas as pd dct = {'col1': 'val1', 'col2': 2} df = pd.DataFrame(dct) ``` So we get the error above: ``` ValueError: If using all scalar values, you must pass an index ``` ## Cause The error is raised because we pass only scalar values. What is a scalar value? **Definition** **Scalar value** is a single value. So the reason is that we are using a dictionary with single values. ```json { "col1": "val1", "col2": 2 } ``` So using single values rather than lists or arrays will raise the error: `ValueError: If using all scalar values, you must pass an index` ## Solution - use vector values So one solution is to use vector values instead of scalar values. In practice this means turning: ```json { "col1": "val1", "col2": 2 } ``` to ```json { "col1": ["val1"], "col2": [2] } ``` This would resolve the error. We can use dict comprehensions in order to workaround the error like - `{k:[v] for k,v in dct.items()}`: ```python dct = {'col1': 'abc', 'col2': 123} dct = {k:[v] for k,v in dct.items()} pd.DataFrame(dct) ``` This will solve the error and create DataFrame like: | | col1 | col2 | | - | ---- | ---- | | 0 | abc | 123 | ## Solution - wrap with list Another solution is to wrap the whole dictionary with list: ```python dct = [{'col1': 'abc', 'col2': 123}] df = pd.DataFrame(dct) ``` The code above will create DataFrame without error. | | col1 | col2 | | - | ---- | ---- | | 0 | abc | 123 | or: ```python dct = {'col1': ['abc'], 'col2': [123]} df = pd.DataFrame(dct) ``` which has the same output in this case: | | col1 | col2 | | - | ---- | ---- | | 0 | abc | 123 | ## Solution - add index If we don't pass an index, it will raise an error that the 'DataFrame()' constructor requires an index to be passed. So the final solution is to add a index in order to solve: ``` ValueError: If using all scalar values, you must pass an index ``` ```python dct = {'col1': 'abc', 'col2': 123} pd.DataFrame(dct, index=[0]) ``` ## Conclusion We've explained Pandas error **"ValueError: If using all scalar values, you must pass an index"**. Then, we discussed how to produce the error and the cause of the problem. Lastly, we discussed several solutions to resolve the error. Article starts with nice visualization and quick fixes for the error. ### Extract Day, Night, Morning, Afternoon, Evening from Pandas /Python Datetime URL: https://datascientyst.com/extract-day-night-morning-afternoon-evening-from-pandas-python-datetime/ Last updated: 2022-10-11T08:04:04.000Z In this guide, we will see how to extract day, night, morning, afternoon, evening from Pandas DataFrame. We would like to map and return information about part of the day from datetime or string column in DataFrame. Below you can find short answer: **(1) Get Day or Night from datetime** ```python mask = (pd.to_timedelta(df['time']).between(pd.Timedelta('6h'),pd.Timedelta('18h'))) df['new'] = np.where(mask, 'Day', 'Night') ``` **(2) Get morning, noon, evening** ```python df['day_part'] = (pd.to_datetime(df["date"]).dt.hour % 24 + 4) // 4 mapping = {1: 'Late Night', 2: 'Early Morning', 3: 'Morning', 4: 'Noon', 5: 'Evening', 6: 'Night'} df['day_part'].replace(mapping) ``` ## Setup Suppose that we have a dataset which contains the following values with date and time information: ```python import pandas as pd dict = {'date': {0: '28-01-2022 5:25:00 PM', 1: '27-02-2022 6:25:00 PM', 2: '30-03-2022 7:25:00 PM', 3: '29-04-2022 8:25:00 PM', 4: '31-05-2022 9:25:00 PM'}, 'time': {0: '5:25:00', 1: '6:25:00', 2: '7:25:00', 3: '8:25:00', 4: '9:25:00'}} df = pd.DataFrame(dict) ``` data: | | date | time | | - | --------------------- | ------- | | 0 | 28-01-2022 5:25:00 PM | 5:25:00 | | 1 | 27-02-2022 6:25:00 PM | 6:25:00 | | 2 | 30-03-2022 7:25:00 PM | 7:25:00 | | 3 | 29-04-2022 8:25:00 PM | 8:25:00 | | 4 | 31-05-2022 9:25:00 PM | 9:25:00 | Let's check how to extract part of the day from this DataFrame. If you have datetime stored as a string you can convert to to datetime by: ```python df["date"] = pd.to_datetime(df["date"]). ``` ## Step 1: Extract day or night If we like to extract day or night we can use `pd.to_timedelta()` to calculate the difference of the hours. Then we can map as follows: - 6 - 18 - Day - 0 - 6; 18 - 0 - Night ```python mask = (pd.to_timedelta(df['time']).between(pd.Timedelta('6h'),pd.Timedelta('18h'))) df['day_night'] = np.where(mask, 'Day', 'Night') df['day_night'] ``` So the result is: ``` 0 Night 1 Day 2 Day 3 Day 4 Day Name: day_night, dtype: object ``` How does it work? First we calculate the timedelta by: ```python pd.to_timedelta(df['time']) ``` which give us: ``` 0 0 days 05:25:00 1 0 days 06:25:00 2 0 days 07:25:00 3 0 days 08:25:00 4 0 days 09:25:00 Name: time, dtype: timedelta64[ns] ``` Then use `.between()` to find if it is between 6 or 18 hours. Finally we use the mask to map day and night. ## Step 2: Extract morning, noon, evening or night What if we need to extract more parts of daytime like: - early morning - morning - noon - evening - night - late night If we divide the day night cycle into 6 equal parts we can do: ```python df['day_part'] = (pd.to_datetime(df["date"]).dt.hour % 24 + 4) // 4 mapping = {1: 'Late Night', 2: 'Early Morning', 3: 'Morning', 4: 'Noon', 5: 'Evening', 6: 'Night'} df['day_part'].replace(mapping) ``` This would give use: ``` 0 Evening 1 Evening 2 Evening 3 Night 4 Night Name: day_part, dtype: object ``` This solution is taken from: [Get part of day (morning, afternoon, evening, night) in Python dataframe](https://stackoverflow.com/a/59577864?ref=datascientyst.com) ## Step 3: Custom mapping of day parts What if we need to use custom mapping - for example different number of day parts or duration. Then we can build function and apply it to Pandas column as follows: ```python def get_part_of_day(h): if type(h) == int: return ( "morning" if 5 <= h <= 11 else "afternoon" if 12 <= h <= 17 else "evening" if 18 <= h <= 22 else "night" ) else: 'error' df['h'] = pd.to_datetime(df["date"]).dt.hour df['h'].apply(get_part_of_day) ``` First we get the hour for each record. Then we map result to: - 5 - 11 is morning - 12 - 17 - afternoon - 18 - 22 - evening - all the rest is night This mapping is taken from: [part\_of\_day.py](https://gist.github.com/rbw/cd9ce89398a08e2f6b2acab14e3c58d7?ref=datascientyst.com) ![](https://datascientyst.com/content/images/2022/10/extract-day-night-morning-afternoon-evening-from-pandas-python-datetime.png) ## Conclusion This post gave several solutions on how to extract day or night, morning, noon or evening. We saw how to use custom mapping to day parts or a generic mapping. We saw how to map date and time in Python and Pandas to different day periods. ### ValueError: Mixing dicts with non-Series may lead to ambiguous ordering - Pandas URL: https://datascientyst.com/valueerror-mixing-dicts-with-non-series-may-lead-to-ambiguous-ordering-pandas/ Last updated: 2022-10-10T05:22:47.000Z In this tutorial, we'll see how to **solve a Pandas error – "ValueError: Mixing dicts with non-Series may lead to ambiguous ordering."**. We get this error from the Pandas when we try to create DataFrame with mixed elements: - dictionaries - non-Series - list - etc In short to solve this error use `json_normalize()`: ```python from pandas import json_normalize json_normalize(data) ``` ## Reproduce the error First let's see an example with this error. Suppose we have data for the Game of Thrones movie - data is extracted from imdb by library - [cinemagoer](https://pypi.org/project/cinemagoer/?ref=datascientyst.com). If you like to find how to extract and analyze IMDB with Python you can follow youtube channel - [DataScientYst](https://www.youtube.com/channel/UC1KTMbPbmesKSA05Nr8pdKg?ref=datascientyst.com) \- we are planning video on this topic. We would like to create DataFrame with this data like: ```python import pandas as pd data = {'title': 'Game of Thrones', 'year': 2011, 'kind': 'tv series', 'taglines': ['Winter is coming.', 'Winter is here. (season 7)', 'The Great War Is Here (Season 8)', 'For the Throne.'], 'number of votes': {10: 1210366, 9: 444762, 8: 192847, 7: 76850, 6: 30220, 5: 18478, 4: 9341, 3: 7930, 2: 7145, 1: 65485}, 'arithmetic mean': 9.0, 'median': 1} pd.DataFrame(data) ``` we got error like: ``` ValueError: Mixing dicts with non-Series may lead to ambiguous ordering. ``` the one mentioned in the title. ## Solve the error - json\_normalize To solve this error we will use: `json_normalize` ```python from pandas import json_normalize json_normalize(data) ``` Now we can create DataFrame from the input data without error. DataFrame below is transposed - for readability: | | 0 | | ------------------ | ---------------------------------------------------------------------------------------------------- | | title | Game of Thrones | | year | 2011 | | kind | tv series | | taglines | \[Winter is coming., Winter is here. (season 7), The Great War Is Here (Season 8), For the Throne.\] | | arithmetic mean | 9.0 | | median | 1 | | number of votes.10 | 1210366 | | number of votes.9 | 444762 | | number of votes.8 | 192847 | | number of votes.7 | 76850 | | number of votes.6 | 30220 | | number of votes.5 | 18478 | | number of votes.4 | 9341 | | number of votes.3 | 7930 | | number of votes.2 | 7145 | | number of votes.1 | 65485 | ## Solve the error - json Alternatively we can read only the important information for us by: ```python import pandas as pd df = pd.DataFrame(data["taglines"]) ``` which give us: | | 0 | | - | -------------------------------- | | 0 | Winter is coming. | | 1 | Winter is here. (season 7) | | 2 | The Great War Is Here (Season 8) | | 3 | For the Throne. | or if we read json file we can use Python json library to load the file as: ```python import json data = json.load(open('data.json')) df = pd.DataFrame(data["taglines"]) ``` ### Convert Pandas to Dask DataFrame ( Dask to Pandas ) URL: https://datascientyst.com/convert-pandas-to-dask-dataframe-dask-to-pandas/ Last updated: 2022-10-09T06:10:44.000Z In this short article, we will see how to **convert Pandas DataFrame to Dask DataFrame.** We will also cover conversion from Dask to Pandas DataFrame. **(1) Convert Pandas to Dask DataFrame** ```python from dask import dataframe as dd df_dd = dd.from_pandas(df, npartitions=2) ``` **(2) Convert Dask to Pandas DataFrame** ```python df_dd.compute() ``` Let's cover both cases in examples and more details. ## Pandas vs Dask DataFrame First let's start with few words about **difference between Pandas and Dask DataFrames.** There is nice reading from Tom Augspurger [Modern Pandas (Part 8): Scaling](http://tomaugspurger.github.io/modern-8-scaling.html?ref=datascientyst.com). From this article we can read: > You can't have a DataFrame larger than your machine's RAM. In practice, your available RAM should be several times the size of your dataset So one of the problems with Pandas is the large datasets. On the other hand, Dask allows you to have a DataFrame larger than your RAM. And another important information: > A Dask DataFrame consists of many pandas DataFrames arranged by the index. Dask is really just coordinating these pandas DataFrames. So **Dask DataFrame consists of Pandas DataFrames arranged by the index.** ![](https://datascientyst.com/content/images/2022/10/convert-pandas-to-dask-dataframe-dask-to-pandas.png) Use Pandas when your machine has enough memory. Otherwise use Dask DataFrame and if needed extract a smaller subset from the Dask DataFrame. Both share very common API - if you like to find how to work with Dask please read: [How to Read and Analyze a Large CSV File With Pandas/Dask](https://datascientyst.com/read-analyse-large-csv-file-pandas-dask/) ## 1: Convert Pandas to Dask DataFrame First we will see how to Convert Pandas to Dask DataFrame. Suppose we have Pandas DataFrame as the one below: ```python import pandas as pd data = {'A':[11,12,13],'B':[21,22,23],'C':[31,32,33]} df = pd.DataFrame(data) ``` Which looks like: | | A | B | C | | - | -- | -- | -- | | 0 | 11 | 21 | 31 | | 1 | 12 | 22 | 32 | | 2 | 13 | 23 | 33 | ### 1.1 Convert Pandas to Dask To **convert it to Dask DataFrame** we can use Dask method `.from_pandas()`: ```python from dask import dataframe as dd df_dd = dd.from_pandas(df, npartitions=3) ``` This give us: ``` Dask DataFrame Structure: A B C npartitions=2 0 int64 int64 int64 1 ... ... ... 2 ... ... ... Dask Name: from_pandas, 2 tasks ``` We can see the types of both DataFrames by: ```python type(df_dd) type(df) ``` and result is: ``` dask.dataframe.core.DataFrame pandas.core.frame.DataFrame ``` ### 1.2 dd.from\_pandas() parameters The method has signature: ```python dd.from_pandas( data, npartitions=None, chunksize=None, sort=True, name=None, ) ``` where important are 2 parameters: - `npartitions` \- number of partitions - `chunksize` \- size of each partition ### 1.3 ValueError: Exactly one of npartitions and chunksize must be specified For example error like: > ValueError: Exactly one of npartitions and chunksize must be specified. is raised when on of: - `npartitions` - `chunksize` is missing. So at this point when you convert you need to know: - how many partitions (sub Pandas DataFrames ) there will be - what is the size for those sub Pandas DataFrames which makes the Dask DataFrame The answer depends on: - system memory - original dataset/DataFrame memory - what operations will be done on the Dask DataFrame ## 2: Convert Dask to Pandas DataFrame To convert from Dask to Pandas DataFrame we can use method `.compute()`: ```python df_pd = df_dd.compute() ``` Now we get again: | | A | B | C | | - | -- | -- | -- | | 0 | 11 | 21 | 31 | | 1 | 12 | 22 | 32 | | 2 | 13 | 23 | 33 | And the types are: ```python type(df_dd) type(df_pd) ``` and result is: ``` dask.dataframe.core.DataFrame pandas.core.frame.DataFrame ``` ### Pandas Exercises - View and Explore Data - Part 1 URL: https://datascientyst.com/pandas-exercises-view-explore-data-solutions/ Last updated: 2022-10-08T06:53:55.000Z This is Pandas exercise about data analysis and exploration. It's the first exercise from a series (check the end of the post for more details). You can test knowledge in several areas: - initial data analysis - "viewing data" - as advertised on the Pandas page: [Viewing data](https://pandas.pydata.org/docs/user%5Fguide/10min.html?ref=datascientyst.com#viewing-data) - summarizing data We can think of data exploration as getting to know your food. I think that most of us prefer to see, taste and smell the food - before eating. For me it's similar to data. Data should be "seen, smelled and tasted". ## Setup In this tutorial we will work with the following dataset: [The Movies Dataset](https://www.kaggle.com/datasets/rounakbanik/the-movies-dataset/versions/5?resource=download&select=movies%5Fmetadata.csv&ref=datascientyst.com). The dataset contains data for 45,000 movies. Data points include cast, crew, plot keywords, revenue, release dates, languages, production companies, countries, vote averages and more. We will focus on the file: `movies_metadata.csv`. Data from the CSV file can be read as follows: ```python import pandas as pd df = pd.read_csv('movies_metadata.csv', low_memory=False) ``` ## 1: How many rows and columns? Usually the first thing which I do is to get the DataFrame dimension. It will give us a rough idea about the size of the DataFrame. Why rough? Because you may have nested hierarchical data - multilevel columns. How can we see the number of rows and columns in Pandas? ```python df.shape ``` This will give us: ``` (45466, 24) ``` This might save you time and effort. Sometimes it is good to start with smaller data samples. ### Bonus question 1 Can you guess how many rows and columns? So to get number of rows we can do: ```python df.shape[0] ``` ### Bonus Question 2 Can you name synonyms or other names of columns and rows? Synonyms of rows and columns **rows** \- observations, records, trials **columns** \- variable, feature ## 2: View the top and bottom rows How can we see the first N records of the dataset? What about the last N rows? Display first and last rows of DataFrame \- top rows ```python df.head() ``` \- last rows ```python df.tail() ``` ### Bonus question 3 How can we see bottom and top rows simultaneously in Pandas? Top and bottom rows simultaneously ```python import numpy as np df.iloc[np.r_[0:5, -5:0]] ``` Did you solve that one? Good job, it was a tricky one! ### Bonus question 4 Do you know why there are parentheses and sometimes not? like: ```python df.head() ``` and: ```python df.shape ``` Parentheses in Pandas? The reason is because ```python head() ``` is a method and should be with parentheses. While ```python shape ``` is an attribute - so it should be without parentheses ## 3: Display the index and columns Can you display the column names? What about index/row names? List column/index names and size ```python df.columns ``` ```python df.index ``` We will get about the columns: ``` Index(['adult', 'belongs_to_collection', 'budget', 'genres', 'homepage', 'id', 'imdb_id', 'original_language', 'original_title', 'overview', 'popularity', 'poster_path', 'production_companies', 'production_countries', 'release_date', 'revenue', 'runtime', 'spoken_languages', 'status', 'tagline', 'title', 'video', 'vote_average', 'vote_count'], dtype='object') ``` and for the index: ``` RangeIndex(start=0, stop=45466, step=1) ``` ## 4: Generate descriptive statistics What about getting a quick statistical summary for a DataFrame in Pandas? This would give us valuable information about: - count - mean - standard deviation - percentile To get summary stats in Pandas we can use method: ```python df.describe() ``` This would give us: | | revenue | runtime | vote\_average | vote\_count | | ----- | ------------ | ------------ | ------------- | ------------ | | count | 4.546000e+04 | 45203.000000 | 45460.000000 | 45460.000000 | | mean | 1.120935e+07 | 94.128199 | 5.618207 | 109.897338 | | std | 6.433225e+07 | 38.407810 | 1.924216 | 491.310374 | | min | 0.000000e+00 | 0.000000 | 0.000000 | 0.000000 | | 25% | 0.000000e+00 | 85.000000 | 5.000000 | 3.000000 | | 50% | 0.000000e+00 | 95.000000 | 6.000000 | 10.000000 | | 75% | 0.000000e+00 | 107.000000 | 6.800000 | 34.000000 | | max | 2.787965e+09 | 1256.000000 | 10.000000 | 14075.000000 | The method will analyze both numeric and object series, as well as DataFrame column sets of mixed data types. ## 5: What are the dtypes? What are the data types of the columns in a DataFrame. Data types might be several kinds like: - numeric - datetime - string/object - nested data - dictionary, list or json - categorical It can help us when we work with our data. Think of it like knowing the OS of the computer. Working on Linux differs from Windows? At least for me. So how to get data types? Show data types of DataFrame ```python df.dtypes ``` This would give us: ``` adult object belongs_to_collection object budget object genres object homepage object id object imdb_id object original_language object original_title object overview object popularity object poster_path object production_companies object production_countries object release_date object revenue float64 runtime float64 spoken_languages object status object tagline object title object video object vote_average float64 vote_count float64 dtype: object ``` ## 6: Transpose rows How can we see the first N records of the dataset - transposed? I like to transpose rows when I'm working with few rows with many columns. The first few rows looks like that(transposed): | | 0 | 1 | | ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | adult | False | False | | belongs\_to\_collection | {'id': 10194, 'name': 'Toy Story Collection', 'poster\_path': '/7G9915LfUQ2lVfwMEEhDsn3kT4B.jpg', 'backdrop\_path': '/9FBwqcd9IRruEDUrTdcaafOMKUq.jpg'} | NaN | | budget | 30000000 | 65000000 | | genres | \[{'id': 16, 'name': 'Animation'}, {'id': 35, 'name': 'Comedy'}, {'id': 10751, 'name': 'Family'}\] | \[{'id': 12, 'name': 'Adventure'}, {'id': 14, 'name': 'Fantasy'}, {'id': 10751, 'name': 'Family'}\] | | homepage | http://toystory.disney.com/toy-story | NaN | | id | 862 | 8844 | | imdb\_id | tt0114709 | tt0113497 | | original\_language | en | en | | original\_title | Toy Story | Jumanji | | overview | Led by Woody, Andy's toys live happily in his room until Andy's birthday brings Buzz Lightyear onto the scene. Afraid of losing his place in Andy's heart, Woody plots against Buzz. But when circumstances separate Buzz and Woody from their owner, the duo eventually learns to put aside their differences. | When siblings Judy and Peter discover an enchanted board game that opens the door to a magical world, they unwittingly invite Alan -- an adult who's been trapped inside the game for 26 years -- into their living room. Alan's only hope for freedom is to finish the game, which proves risky as all three find themselves running from giant rhinoceroses, evil monkeys and other terrifying creatures. | | popularity | 21.946943 | 17.015539 | | poster\_path | /rhIRbceoE9lR4veEXuwCC2wARtG.jpg | /vzmL6fP7aPKNKPRTFnZmiUfciyV.jpg | | production\_companies | \[{'name': 'Pixar Animation Studios', 'id': 3}\] | \[{'name': 'TriStar Pictures', 'id': 559}, {'name': 'Teitler Film', 'id': 2550}, {'name': 'Interscope Communications', 'id': 10201}\] | | production\_countries | \[{'iso\_3166\_1': 'US', 'name': 'United States of America'}\] | \[{'iso\_3166\_1': 'US', 'name': 'United States of America'}\] | | release\_date | 1995-10-30 | 1995-12-15 | | revenue | 373554033.0 | 262797249.0 | | runtime | 81.0 | 104.0 | | spoken\_languages | \[{'iso\_639\_1': 'en', 'name': 'English'}\] | \[{'iso\_639\_1': 'en', 'name': 'English'}, {'iso\_639\_1': 'fr', 'name': 'Français'}\] | | status | Released | Released | | tagline | NaN | Roll the dice and unleash the excitement! | | title | Toy Story | Jumanji | | video | False | False | | vote\_average | 7.7 | 6.9 | | vote\_count | 5415.0 | 2413.0 | And transposing rows is possible by: Transpose rows to columns ```python df.head(2).T ``` Which is the alias of method `transpose()`. This is another thing to have in mind working with Pandas - aliases. ## 7: Sort values Can you get the top most rated movies from this dataset? To do so we will need to sort data by a column. So the last exercise is to get the names and number of votes for the most voted movies in that dataset. Find 5 top voted movies ```python df.sort_values(by='vote_count', ascending=False)[['title', 'vote_count']].head() ``` ### Bonus question 5 Can you find the top 7 values for the numerical columns? Return the result as a DataFrame: | | revenue | runtime | vote\_average | vote\_count | | - | ------------ | ------- | ------------- | ----------- | | 0 | 2.787965e+09 | 1256.0 | 10.0 | 14075.0 | | 1 | 2.068224e+09 | 1140.0 | 10.0 | 12269.0 | | 2 | 1.845034e+09 | 1140.0 | 10.0 | 12114.0 | | 3 | 1.519558e+09 | 931.0 | 10.0 | 12000.0 | | 4 | 1.513529e+09 | 925.0 | 10.0 | 11444.0 | | 5 | 1.506249e+09 | 900.0 | 10.0 | 11187.0 | | 6 | 1.405404e+09 | 877.0 | 10.0 | 10297.0 | Can you optimize the proposed solution below? Find top values for numerical columns ```python from pandas.api.types import is_numeric_dtype dfs = [] for col in df.columns: top_values = [] if is_numeric_dtype(df[col]): top_values = df[col].nlargest(n=7) dfs.append(pd.DataFrame({col: top_values}).reset_index(drop=True)) pd.concat(dfs, axis=1) ``` Were you able to solve it? Well done! ### Bonus question 6 Method `describe()` is good for initial exploratory data analysis. For more details like: - type - unique values - missing values - most frequent values - histogram and more Can you research and find Python library which can do "serious exploratory data analysis"? Open Collapsible We can use library: [pandas-profiling · PyPI](https://pypi.org/project/pandas-profiling/?ref=datascientyst.com) ```python ProfileReport(df, title="Pandas Profiling Report") ``` ## Conclusion Congratulations! You have passed the first set of Pandas exercises. We've tasted the movie dataset. Now you will have a better understanding of your data and how to do initial data analysis. In this post, we learned about initial data analysis in Pandas. Most popular methods in Pandas for "viewing data". We saw some advanced techniques and tricks very useful when data is touched for the first time. ### Pandas exercises for beginners Pandas exercises will be structured as a sample project plus exercises - a place where we can practice data science. The goal is to test your knowledge and learn in 3 main areas: - data cleaning - data analysis - data collection The areas above are often reported as problematic in data science. Data preparation is blamed to consume up to 80% from Data science projects. Dirty data is one of the main factors for failed projects in data science and machine learning. ### How to Read and Analyse a Large CSV File With Pandas/Dask URL: https://datascientyst.com/read-analyse-large-csv-file-pandas-dask/ Last updated: 2023-03-17T23:04:43.000Z In this article, we will see **how to read and analyse large CSV or text files with Panda**s. This can be used when your machine doesn't have enough memory to process the whole file. So you will learn at the end how to use Pandas for big data. In short you can try by using Dask which is a wrapper of Pandas: ```python import dask.dataframe as dd df = dd.read_csv('huge_file.csv') ``` ## Setup Often genome data has huge files often more than 30 GB. To read such files from a laptop or machine with less than 30 GB we will need to use a library like Dask. To get sample huge files for tests we can get from: - Genome data - [https://www.sanger.ac.uk/resources/downloads/human/](https://www.sanger.ac.uk/resources/downloads/human/?ref=datascientyst.com) - Kaggle filters - [https://www.kaggle.com/datasets?fileType=csv&sizeStart=30%2CGB](https://www.kaggle.com/datasets?fileType=csv&sizeStart=30%2CGB&ref=datascientyst.com) To download such file we can use: ```python wget ftp://ftp.ensembl.org/pub/release-77/fasta/homo_sapiens/dna/Homo_sapiens.GRCh38.dna.toplevel.fa.gz gunzip Homo_sapiens.GRCh38.dna.toplevel.fa.gz ``` The file will download as an archive. After uncompressing the file size will be 37GB. ## Install Dask First we need to install Dask library: ```bash pip install dask ``` To read more about this library from the official documentation we can visit: [Get Started with Dask](https://www.dask.org/get-started?ref=datascientyst.com) Benefits of using Dask are: - similar API like Pandas - cluster scaling - parallel computations - read "big data files" - Python compatible You can read more here: [Why Dask?](https://docs.dask.org/en/stable/why.html?ref=datascientyst.com) ## Read a Large CSV File To read large CSV file with Dask in Pandas similar way we can do: ```python import dask.dataframe as dd df = dd.read_csv('huge_file.csv') ``` We can also read archived files directly without uncompression but often there are problems. So when possible try to uncompress the file before reading it. ## Work with Dask DataFrame Some operation like: ```python df.head() ``` will work fast and without errors. Other like: ```python df.tail() ``` may lead to errors: > ValueError: Mismatched dtypes found in `pd.read_csv`/`pd.read_table`. Check the final section in this article to solve them. Other will result into: ```python df.shape ``` into delayed operations: ``` (Delayed('int-af70a7d8-597c-4864-a876-2d74003f0e97'), 125) ``` ## Get shape of Dask DataFrame To find number of the rows of a Dask DataFrame or the `df.shape` we can use: ```python t = df.shape t[0].compute(),t[1] ``` this requires scan of whole data range: result: ``` (622038697, 1) ``` ## Initial investigation on Dask DataFrame To start analyzing Dask DataFrame we can start by: ```python df.head() ``` which will work normally as in Pandas. We can also use it for other operation in order to analyze higher number of rows: ```python df.head(1000000).iloc[:, 0].str.find('AAACCCAAA') ``` this results into: ``` 0 -1 1 -1 2 -1 3 -1 4 -1 .. 999995 -1 999996 -1 999997 -1 999998 -1 999999 -1 Name: >1 dna:chromosome chromosome:GRCh38:1:1:248956422:1 REF, Length: 1000000, dtype: int64 ``` If we try the same - search the first column of DataFrame for string pattern - for the whole DataFrame we will got: ```python df.iloc[:, 0].str.find('AAACCCAAA') ``` result: ``` Dask Series Structure: npartitions=593 int64 ... ... ... ... Name: >1 dna:chromosome chromosome:GRCh38:1:1:248956422:1 REF, dtype: int64 Dask Name: str-find, 1779 tasks ``` Information about the partitions and number of tasks but not the actual result. ## Dask .compute() To get results from Dask operations we use `.compute()`. Let's check how to search the whole Dask DataFrame for a given string: ```python t = df.iloc[:, 0].str.find('AAACCCAAA') t.compute() ``` Now we will get results for the whole DataFrame if string 'AAACCCAAA' is part of the first column or not: ``` 0 -1 1 -1 2 -1 3 -1 4 -1 .. 923939 -1 923940 -1 923941 -1 923942 -1 923943 -1 Name: >1 dna:chromosome chromosome:GRCh38:1:1:248956422:1 REF, Length: 622038697, dtype: int64 ``` Note that those operations can take longer time due to the size, memory required and number of tasks. ## Dask show progress bar Finally let's cover how to show progress bar for the operations of Dask. To do so we can use `tqdm` or Dask diagnostic - ProgressBar. There is dask integration for it: `from dask.diagnostics import ProgressBar`. So suppose that we have: ```python t = df.iloc[:, 0].str.find('AAACCCAAA') ``` ### ProgressBar So to get progress bar while working with big CSV files we can do: ```python from dask.diagnostics import ProgressBar from dask import delayed,compute with ProgressBar(): compute(t) ``` This would show progress bar like: ``` [### ] | 8% Completed | 1min 43.7s [###### ] | 15% Completed | 2min 35.8s [########## ] | 25% Completed | 3min 55.6s ``` ### tqdm So with tqdm we can try something like: ```python from tqdm.dask import TqdmCallback from dask.diagnostics import ProgressBar ProgressBar().register() t = df.iloc[:, 0].str.find('AAACCCAAA') with TqdmCallback(desc="compute"): t.compute() ``` result: ``` [ ] | 0% Completed | 0.0s compute: 0%| | 0/593 [00:00 8]['Time'] ``` Which will give us: ``` 3378 1975-02-23T02:58:41.000Z 7512 1985-04-28T02:53:41.530Z 20650 2011-03-13T02:23:34.520Z Name: Time, dtype: object ``` To exclude rows with different rows we can do: ```python df = df[df['Date'].str.len() == 8] ``` This way is better for for working with date or time formats like: - HH:MM:SS - dd/mm/YYYY ## Infer date and time format from datetime This option is best when we need to work with formats which has date and time like: - '%Y-%m-%dT%H:%M:%S.%f' - '%Y-%m-%dT%H:%M:%S' First we will find the most frequent format in the column by: ```python import numpy as np from pandas.core.tools.datetimes import _guess_datetime_format_for_array array = np.array(df["Date"].to_list()) _guess_datetime_format_for_array(array) ``` This would give us: ``` '%m/%d/%Y' ``` Next we will try to convert the whole column with this format: ```python pd.to_datetime(df["Date"], format='%m/%d/%Y') ``` This will result into error: ``` ValueError: time data '1975-02-23T02:58:41.000Z' does not match format '%m/%d/%Y' (match) ``` Now we can exclude rows with this format or convert them with different format: ## More errors related to pd.to\_datetime() ### Semantic date errors in Pandas Sometimes there isn't a code error but the date is wrong. This is when there are parsing errors. For example day and month are wrongly inferred or there are two date formats. To find more to this problem check: [How to Fix Pandas to\_datetime: Wrong Date and Errors](https://datascientyst.com/how-to-fix-pandas-to%5Fdatetime-wrong-date-and-errors/) ### Convert string to datetime If you want to find more about convert string to datetime and infer date formats you can check: [Convert String to DateTime in Pandas](https://datascientyst.com/convert-string-to-datetime-pandas/) ### How to Extract Hour and Minute in Pandas URL: https://datascientyst.com/how-to-extract-hour-and-minute-in-pandas/ Last updated: 2022-10-05T07:31:25.000Z In this tutorial, we're going to **extract hours and minutes in Pandas**. We will extract information from time stored as string or datetime. Here are several ways to extract hours and minutes in Pandas: **(1) Use accessor on datetime column** ```python df['date'].dt.hour ``` **(2) Apply lambda and datetime** ```python df['date'].apply(lambda x: datetime.strptime(x, '%d-%m-%Y %H:%M:%S %p').hour) ``` ## Setup Let's work with the following DataFrame which has time information stored as a string: ```python import pandas as pd dict = {'date': {0: '28-01-2022 5:25:00 PM', 1: '27-02-2022 6:25:00 PM', 2: '30-03-2022 7:25:00 PM', 3: '29-04-2022 8:25:00 PM', 4: '31-05-2022 9:25:00 PM'}, 'date_short': {0: 'Jan-2022', 1: 'Feb-2022', 2: 'Mar-2022', 3: 'Apr-2022', 4: 'May-2022'}} df = pd.DataFrame(dict) ``` data: | | date | date\_short | | - | --------------------- | ----------- | | 0 | 28-01-2022 5:25:00 PM | Jan-2022 | | 1 | 27-02-2022 6:25:00 PM | Feb-2022 | | 2 | 30-03-2022 7:25:00 PM | Mar-2022 | | 3 | 29-04-2022 8:25:00 PM | Apr-2022 | | 4 | 31-05-2022 9:25:00 PM | May-2022 | We can't extract hour and minutes reliably just by using string operations. So in the next steps we will see how to extract time info in a reliable way. ## Step 1: Convert string to time First we will convert the string column to datetime. To do so in Pandas we can use method `to_datetime(df['date'])`: ```python pd.to_datetime(df['date']) ``` result: ``` 0 2022-01-28 17:25:00 1 2022-02-27 18:25:00 2 2022-03-30 19:25:00 3 2022-04-29 20:25:00 4 2022-05-31 21:25:00 Name: date, dtype: datetime64[ns] ``` ## Step 2: Extract hours from datetime To extract hours from datetime or timestamp in Pandas we can use accessors: `.dt.hour`. So working with datetime column we get: ```python df['date'].dt.hour ``` Hours represented in 24 hour format: ``` 0 17 1 18 2 19 3 20 4 21 Name: date, dtype: int64 ``` Because we had initial information stored as `PM` The image below shows the extract hours: ![](https://datascientyst.com/content/images/2022/10/how-to-extract-hour-and-minute-in-pandas.png) ## Step 3: Extract minutes from datetime Extracting minutes in Pandas from datetime is available by accessor - `.dt.minute`: ```python df['date'].dt.minute ``` We get minutes from the datetime: ``` 0 25 1 25 2 25 3 25 4 25 Name: date, dtype: int64 ``` ## Step 4: Extract seconds from datetime There is one last option - for seconds - `.dt.second`: ```python df['date'].dt.second ``` We extracted seconds for all rows: ``` 0 0 1 0 2 0 3 0 4 0 Name: date, dtype: int64 ``` ## Step 5: Extract hour in Pandas with lambda We can use Python's datetime to extract hours directly from Pandas columns. To extract hour or minutes we can: - apply `lambda` - select date and time format: ```python from datetime import datetime df['date'].apply(lambda x: datetime.strptime(x, '%d-%m-%Y %H:%M:%S %p').hour) ``` Extract hours and minutes: ``` 0 5 1 6 2 7 3 8 4 9 Name: date, dtype: int64 ``` Refer to: [Infer date format from string](https://datascientyst.com/convert-string-to-datetime-pandas/#step-4-infer-date-format-from-string) to find how to extract time format from string in Pandas If you need to find more about datetime errors in Pandas check: [How to Fix Pandas to\_datetime: Wrong Date and Errors](https://datascientyst.com/how-to-fix-pandas-to%5Fdatetime-wrong-date-and-errors/) ## Conclusion In this article, we saw several ways to extract time information in Pandas. We saw how to **extract hour and minute from datetime and timestamp in Pandas.** ### Exploratory Data Analysis Python and Pandas with Examples URL: https://datascientyst.com/exploratory-data-analysis-pandas-examples/ Last updated: 2024-01-20T02:00:54.000Z This article is about **Exploratory Data Analysis(EDA) in Pandas and Python.** The article will explain step by step how to do Exploratory Data Analysis plus examples. EDA is an important step in Data Science. The goal of EDA is to identify errors, insights, relations, outliers and more. The image below illustrate the data science workflow and where EDA is located: ![Data_visualization_process_v1](https://datascientyst.com/content/images/2022/09/Data_visualization_process_v1.png) Source: [Exploratory Data Analysis - wikipedia](https://en.wikipedia.org/wiki/Exploratory%5Fdata%5Fanalysis?ref=datascientyst.com) ## Background Story Imagine that you are expecting royal guests for dinner. You are asked to research a special menu from a cooking book with thousands of recipes. As they are very pretentious you need to avoid some ingredients or find exact quantities for others. Dinner and launch menus are needed. Unfortunately some recipes are wrong and others are incomplete. How to start? You may start with the main ingredients. Another option is by selecting the meal type and course. There should be a special menu for Vegetarians. The story above reminds of **Exploratory Data Analysis(EDA)**. All the steps and questions are part of understanding unknown territory. Let's try to solve this with Data Science. ![exploratory-data-analysis-python-pandas.opti.webp](https://datascientyst.com/content/images/2024/01/exploratory-data-analysis-python-pandas.opti.webp) ## What is Exploratory Data Analysis? **Exploratory data analysis is like detective work: searching for insights that identify problems and hidden patterns.** Start with one variable at a time, then explore two variables, and so on. **Definition** **Exploratory Data Analysis(EDA)** \- analyze and investigate datasets and summarize their main characteristics and apply visualization methods. ### Tools for EDA: **Most popular Tools:** - Pandas/Python - R ### Techniques for EDA: Visual techniques - [Box plot](https://en.wikipedia.org/wiki/Box%5Fplot?ref=datascientyst.com) - [Histogram](https://en.wikipedia.org/wiki/Histogram?ref=datascientyst.com) - [Heat map](https://en.wikipedia.org/wiki/Heat%5Fmap?ref=datascientyst.com) - [Scatter plot](https://en.wikipedia.org/wiki/Scatter%5Fplot?ref=datascientyst.com) Dimensionality reduction - [Principal component analysis](https://en.wikipedia.org/wiki/Principal%5Fcomponent%5Fanalysis?ref=datascientyst.com) (PCA) ## Exploratory Data Analysis - Example Project Let's do an example on Exploratory Data Analysis using [Food Recipes](https://www.kaggle.com/datasets/sarthak71/food-recipes?ref=datascientyst.com). We will do step by step analysis on this data set and answer on questions like: - What data do we have? - What is the dimension of this data? - Are there any dependent variables? - What are the data types? - Missing data? - Duplicate data? - Correlations? You can find the notebook on and Google Colab: - [GitHub](https://github.com/softhints/Pandas-Exercises-Projects/blob/main/project/Exploratory%20Data%20Analysis.ipynb?ref=datascientyst.com) - [Google Colab](https://colab.research.google.com/github/softhints/Pandas-Exercises-Projects/blob/main/project/Exploratory%20Data%20Analysis.ipynb?ref=datascientyst.com) \- you can play with this notebook without installing Python ## Step 1: Load Data and Initial Analysis First we will import the needed libraries and then load the dataset into Pandas: ```python import pandas as pd import seaborn as sns import matplotlib.pyplot as plt df = pd.read_csv('https://raw.githubusercontent.com/softhints/Pandas-Exercises-Projects/main/data/food_recipes.csv', low_memory=False) ``` ## Step 2: Initial Analysis of Pandas DataFrame We will check the data by using the following methods: - `df` \- returns first and last 5 records; returns number of rows and columns - `head(n)` \- returns first n rows - `tail(n)` \- returns last n rows - `sample(n)` \- sample random n rows The first 2 rows transposed looks like: | | 0 | 1 | | | -------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------- | | recipe\_title | Roasted Peppers And Mushroom Tortilla Pizza Recipe | Thakkali Gotsu Recipe \| Thakkali Curry | Spicy & Tangy Tomato Gravy | | url | https://www.archanaskitchen.com/roasted-peppers-and-mushroom-tortilla-pizza-recipe | https://www.archanaskitchen.com/tomato-gotsu-recipe-spicy-tangy-tomato-gravy | | | record\_health | good | good | | | vote\_count | 434 | 3423 | | | rating | 4.958525 | 4.932223 | | | description | is a quicker version pizza ....: | also known as the is a ... | | | cuisine | Mexican | South Indian Recipes | | | course | Dinner | Lunch | | | diet | Vegetarian | Vegetarian | | | prep\_time | 15 M | 10 M | | | cook\_time | 15 M | 20 M | | | ingredients | Tortillas\|Extra Virgin Olive Oil|Garlic|Mozzarella cheese|Red Yellow or Green Bell Pepper (Capsicum)|Onions|Kalmatta olives|Button mushrooms | Sesame (Gingelly) Oil\|Mustard seeds (Rai/ Kadugu)|Curry leaves|Garlic|Pearl onions (Sambar Onions)|Tomatoes|Tamarind|Turmeric powder (Haldi)|Salt|Jaggery | | | instructions | To begin making the Roasted ... | To begin making Tomato Gotsu Recipe/ .... | | | author | Divya Shivaraman | Archana Doshi | | | tags | Party Food Recipes\|Tea Party Recipes|Mushroom Recipes|Fusion Recipes|Tortilla Recipe|Bell Peppers Recipes | Vegetarian Recipes\|Tomato Recipes|South Indian Recipes|Breakfast Recipe Ideas | | | category | Pizza Recipes | Indian Curry Recipes | | We are lucky with: ```python df.sample(2) ``` as we can see missing values and non latin symbols: | | recipe\_title | url | record\_health | vote\_count | rating | description | cuisine | course | diet | prep\_time | cook\_time | ingredients | instructions | author | tags | category | | ---- | ------------------------------------------------------------------ | -------------------------------------------------------------------------- | -------------- | ----------- | -------- | --------------------------------------------------------------------------------------------------- | ------- | --------- | ---------- | ---------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- | ----------------- | ------------------------------------------------------------------------------------------------------- | --------------- | | 2400 | Coconut Chia Pudding Recipe | https://www.archanaskitchen.com/coconut-chia-pudding-recipe | good | 162 | 4.975309 | Coconut Chia Pudding makes a great snack or even ... | NaN | NaN | NaN | NaN | NaN | Chia Seeds\|Tender coconut water|Coconut milk|Orange|Blueberries|Pistachios|Honey|Mint Leaves (Pudina) | To begin making the Coconut Chia Pudding recipe... | Gauravi Vinay | Healthy Recipes\|Breakfast Recipe Ideas|Pudding Recipes|Coconut Recipes | Dessert Recipes | | 1447 | पंजाबी स्टाइल बूंदी कढ़ी रेसिपी - Punjabi Style Boondi Kadhi Recipe | https://www.archanaskitchen.com/punjabi-style-boondi-kadhi-recipe-in-hindi | good | 685 | 4.954745 | एक प्रसिद्ध कढ़ी है जिसे राजस्थान और पंजाब में बनाया जाता है. इसमें मेथी के दाने भी डाले जाते है ... | Punjabi | Side Dish | Vegetarian | 10 M | 30 M | दही\|बेसन|हरी मिर्च|अदरक लहसुन का पेस्ट|लाल मिर्च पाउडर|हल्दी पाउडर|नमक|घी|लॉन्ग|पूरी काली मिर्च|जीरा|मेथी के दाने|हींग|बूंदी | पंजाबी स्टाइल बूंदी कढ़ी रेसिपी बनाने के लिए सबसे पहले एक बाउल में ... | Archana's Kitchen | Side Dish Recipes\|Indian Lunch Recipes|Dahi Recipes (Curd/Yogurt)|Indian Dinner Recipes|Winter Recipes | Kadhi Recipes | **The initial findings** for this dataset are: **Insights** - dataset has header column - separator is ';' - `NaN` represents missing values - 2 columns with nested data - ingredients and tags - if we take closer look we can find values divided by pipes - "|" - for further analysis - data need to be normalized - there are non latin characters in this dataset Let's continue with the analysis and not rely on luck or first impressions. ## Step 3: Dimension and data types of DataFrame Next we will use several Pandas methods in order to find information for the columns, dimensions, nested data and more: ```python df.shape ``` will return the dimension of this dataset: ``` (8009, 16) ``` We can get columns by: ```python df.columns ``` ``` Index(['recipe_title', 'url', 'record_health', 'vote_count', 'rating', 'description', 'cuisine', 'course', 'diet', 'prep_time', 'cook_time', 'ingredients', 'instructions', 'author', 'tags', 'category'], dtype='object') ``` To get information about the data types we can use: ```python df.info() df.dtypes ``` which results into: ``` RangeIndex: 8009 entries, 0 to 8008 Data columns (total 16 columns): Column Non-Null Count Dtype --- ------ -------------- ----- 0 recipe_title 8009 non-null object 1 url 8009 non-null object 2 record_health 8009 non-null object 3 vote_count 8009 non-null int64 4 rating 8009 non-null float64 5 description 7994 non-null object 6 cuisine 7943 non-null object 7 course 7854 non-null object 8 diet 7858 non-null object 9 prep_time 7979 non-null object 10 cook_time 7979 non-null object 11 ingredients 7997 non-null object 12 instructions 8009 non-null object 13 author 8009 non-null object 14 tags 7930 non-null object 15 category 8009 non-null object dtypes: float64(1), int64(1), object(14) memory usage: 1001.2+ KB ``` The above gives use: - 3 data types - `float64(1), int64(1), object(14)` - memory usage: 1001.2+ KB - information for the rows and the columns By using the method `df.shape()` we can find the total number of rows and columns. Which has synonyms: - `row` \- observation, record, trial - `column` \- variable, feature, characteristic Data insights from this step: **Insights:** - dataset contains - 8009 rows and 16 columns - 8009 observations - 16 characteristics - 2 numeric columns - 2 which need conversion - prep\_time and cook\_time - convert string to int - 15 M -> 15 ## Step 4: Statistics - min, max, mean, percentile We can get useful statistics for numeric and non numeric columns of this DataFrame. ### 4.1\. Descriptive statistics for numeric columns For numeric columns we can use: ```python df.describe() ``` which returns stats like: - mean - standard deviation - percentile - count - min and max | | vote\_count | rating | | ----- | ------------ | ----------- | | count | 8009.000000 | 8009.000000 | | mean | 2268.004495 | 4.888621 | | std | 3683.156570 | 0.077467 | | min | 15.000000 | 3.175705 | | 25% | 494.000000 | 4.865031 | | 50% | 1050.000000 | 4.900553 | | 75% | 2487.000000 | 4.930000 | | max | 80628.000000 | 5.000000 | **Insights:** - mean value is close to 75% - no missing values - big difference between max and 75% percentile for `vote_count` \- potential outliers ### 4.2\. Pandas describe() for non numeric columns For the rest of the columns we can use `include='object'`: ```python df.describe(include='object') ``` the result is: | | recipe\_title | url | record\_health | description | cuisine | course | diet | ingredients | | ------ | -------------------------------------------------- | ---------------------------------------------------------------------------------- | -------------- | ---------------------------------------------- | ------- | ------ | ---------- | ------------------------------------ | | count | 8009 | 8009 | 8009 | 7994 | 7943 | 7854 | 7858 | 7997 | | unique | 8009 | 8009 | 1 | 7989 | 77 | 13 | 10 | 7953 | | top | Roasted Peppers And Mushroom Tortilla Pizza Recipe | https://www.archanaskitchen.com/roasted-peppers-and-mushroom-tortilla-pizza-recipe | good | If you like this recipe, try more recipes like | Indian | Lunch | Vegetarian | Lemons\|Sugar|Salt|Red Chilli powder | | freq | 1 | 1 | 8009 | 6 | 1336 | 1925 | 5478 | 3 | Few columns seems to be more interesting than the rest: **Insights:** - 4 categorical columns - `category`, `cuisine`, `diet` and `course` - 1 column with single value - `record_health` ## Step 5: Inspect single column of DataFrame ```python df.diet.nunique() ``` returns the number of unique values in this column: ``` 10 ``` ```python df.diet.unique() ``` will give us all unique values: ``` array(['Vegetarian', 'High Protein Vegetarian', 'Non Vegeterian', 'No Onion No Garlic (Sattvic)', 'High Protein Non Vegetarian', 'Diabetic Friendly', 'Eggetarian', 'Vegan', 'Gluten Free', nan, 'Sugar Free Diet'], dtype=object) ``` ```python df.diet.value_counts(dropna=False) ``` returns the values and their frequency: ``` Vegetarian 5478 High Protein Vegetarian 792 Non Vegeterian 447 Eggetarian 391 Diabetic Friendly 296 High Protein Non Vegetarian 225 NaN 151 No Onion No Garlic (Sattvic) 77 Vegan 71 Gluten Free 66 Sugar Free Diet 15 Name: diet, dtype: int64 ``` **Insights:** - 10 unique values - categorical column - `Vegetarian` \- is the most frequent value - 151 - missing values ## Step 6: Data conversion in Pandas Sometimes data needs to be processed for proper analysis. For example nested data to be expanded, string needs to be converted to number or date. We'll see two examples of data processing required for further data analysis. ### 6.1 Convert String to numeric To analyze columns `cook_time` and `prep_time` as numeric we can convert them by: ```python df['cook_time'] = df['cook_time'].str.replace(' M', '').fillna(0).astype(int) df['prep_time'] = df['prep_time'].str.replace(' M', '').fillna(0).astype(int) ``` Now we can apply visual techniques to analyze those columns for outliers - this is done in section 7.3 of this article. ### 6.2 Expand nested column Next we will expand a column "ingredients" - in order to analyze ingredients for each recipe: ```python df['ingredients'].str.split('|', expand=True) ``` The result is new DataFrame with expanded values from the ingredient column: from: ``` 0 Tortillas|Extra Virgin Olive Oil|Garlic|Mozzar... 1 Sesame (Gingelly) Oil|Mustard seeds (Rai/ Kadu... 2 Extra Virgin Olive Oil|Pineapple|White onion|R... ``` to: | | 0 | 1 | 2 | 3 | 4 | | - | -------------------------- | --------------------------- | --------------- | -------------------------------------------- | ------------------------------------------ | | 0 | Tortillas | Extra Virgin Olive Oil | Garlic | Mozzarella cheese | Red Yellow or Green Bell Pepper (Capsicum) | | 1 | Sesame (Gingelly) Oil | Mustard seeds (Rai/ Kadugu) | Curry leaves | Garlic | Pearl onions (Sambar Onions) | | 2 | Extra Virgin Olive Oil | Pineapple | White onion | Red Yellow and Green Bell Peppers (Capsicum) | Pickled Jalapenos | | 3 | Arhar dal (Split Toor Dal) | Turmeric powder (Haldi) | Salt | Dry Red Chillies | Mustard seeds (Rai/ Kadugu) | | 4 | Rajma (Large Kidney Beans) | Cashew nuts | Sultana Raisins | Asafoetida (hing) | Cumin seeds (Jeera) | To find the most common 'ingredient' in the all recipes we can use the following steps: - split the column on separator "|" - `expand=True` \- will expand data to DataFrame - `melt()` \- transform DataFrame columns to rows - pairs of variable and value - count the values by `.value_counts()` ```python df['ingredients'].str.split('|', expand=True).melt()['value'].value_counts() ``` we found that salt is the most popular ingredient: ``` Salt 5698 Sunflower Oil 2241 Garlic 2227 Turmeric powder (Haldi) 2221 Onion 2012 ... नारीयल पानी 1 नारीयल का दूध 1 सिंघाड़े का आटा 1 Black currant (dried) 1 Corn chips 1 Name: value, Length: 1823, dtype: int64 ``` ## Step 7: Visualization for Data Analysis Finally lets cover a few visualization techniques to find missing values, outliers and correlations. ### 7.1 Missing values The first visualization technique is about finding missing values in Pandas DataFrame. Use library Seaborn with heat map to plot missing values for the whole dataset: - `df.isna()` \- returns new DataFrame with True/False values depending on the missing values in a cell - `cmap = 'Greens'` \- is a color map. For more info check: [Get color palette with matplotlib](https://datascientyst.com/get-list-of-n-different-colors-names-python-pandas/#step-4-get-color-palette-with-matplotlib) ```python sns.heatmap(df.isna(),cmap = 'Greens') ``` result is: ![detect-missing-values-pandas-heatmap-plot](https://datascientyst.com/content/images/2022/09/detect-missing-values-pandas-heatmap-plot.png) Next let's check the same visualization for the last 3000 rows: ```python sns.heatmap(df.tail(3000).isna(),cmap = 'Greens') ``` result: ![exploratory-data-analysis-pandas-examples-missing-values](https://datascientyst.com/content/images/2022/09/exploratory-data-analysis-pandas-examples-missing-values.png) From the result above we can conclude: **Insights:** - data set has missing values in several columns - represented by darker green lines - `cuisine`, `course` and `diet` - plot suggest pattern for missing values of those columns ### 7.2 Correlations Again will use Seaborn's heatmap in order to visualize correlations between two and more columns: ```python sns.heatmap(df.corr(),cmap='Greens',annot=False) ``` The first option is for the whole dataset shows - no correlation between the columns. ![exploratory-data-analysis-pandas-examples-correlation](https://datascientyst.com/content/images/2022/09/exploratory-data-analysis-pandas-examples-correlation.png) ```python sns.heatmap(df[df.diet=='Vegan'].corr(),cmap='Greens',annot=True) ``` While the second one is showing correlations for rows with type "Vegan". We can see a correlation between columns 'prep\_time' and 'cook\_time'. ![exploratory-data-analysis-pandas-examples-correlation-heatmap](https://datascientyst.com/content/images/2022/09/exploratory-data-analysis-pandas-examples-correlation-heatmap.png) **Insights:** - no correlation of numerical columns for the whole DataFrame - correlation for - Vegan recipes - between `prep_time` and `cook_time` - darker color represent positive correlation - lighter - negative - to get values use `annot=True` ### 7.3 Detect outliers in Pandas DataFrame Finally we will use boxplot to detect outliers: ```python cols = ['vote_count', 'rating', 'prep_time', 'cook_time'] plt.figure(figsize=(20, 5)) sns.boxplot(data=df[cols], orient='h') ``` The picture below suggest about outliers in the numeric columns ![exploratory-data-analysis-pandas-examples-outliers](https://datascientyst.com/content/images/2022/09/exploratory-data-analysis-pandas-examples-outliers.png) **Insights:** - `vote_count` \- suggest outliers - `cook_time` \- has possible outliers - `rating` \- doesn't indicate outliers ### 7.4 Explain Seaborn boxplot and outliers In order to get a better understanding we will work only with two columns: `'prep_time', 'cook_time'` and the first 18 records. The stats for them are: ```python df[['prep_time', 'cook_time']].head(18).describe() ``` | | prep\_time | cook\_time | | ----- | ---------- | ---------- | | count | 18.000000 | 18.000000 | | mean | 12.222222 | 22.388889 | | std | 4.608886 | 17.091928 | | min | 5.000000 | 0.000000 | | 25% | 10.000000 | 11.250000 | | 50% | 10.000000 | 20.000000 | | 75% | 15.000000 | 30.000000 | | max | 20.000000 | 60.000000 | The picture below illustrates how to analyze a dataset for outliers with boxplot. ![detecting_outliers_in_python_pandas_boxplot](https://datascientyst.com/content/images/2022/09/detecting_outliers_in_python_pandas_boxplot.png) Boxplot has 3 main components: - box (Interquartile Range - IQR) - q1 - 25% - q2 - 50% or median - q3 - 75% - whiskers - min - max - outliers - points beyond the whiskers We can notice that sometimes the whiskers don't match the min/max. They are calculated based on formulas. Upper and low whiskers are calculated as 1.5 multiplied on the interquartile range - IQR(the box). **Definition** An **outlier(extreme)** is a value that lies in the extremes of a data Series(on either end). They can affect overall observation and insights. In the next article we will cover more techniques and several Python libraries for exploratory data analysis. ## Conclusion In this article, we covered the basics of Exploratory Data Analysis and discussed most of the common methods which every data scientist should know while working with datasets. Through practical examples, we learned how to explore new datasets. We started by initial data analysis and first glance of data. After that, we made data conversion and preprocessing for further analysis. Finally, we saw multiple visual techniques in order to find missing values, outliers and correlations. ### How to Extract Table from PDF with Python and Pandas URL: https://datascientyst.com/extract-table-from-pdf-with-python-pandas/ Last updated: 2025-02-14T13:10:18.000Z In this short tutorial, we'll see how to **extract tables from PDF files with Python and Pandas**. We will cover two cases of table extraction from PDF: **(1) Simple table with tabula-py** ```python from tabula import read_pdf df_temp = read_pdf('china.pdf') ``` **(2) Table with merged cells** ```python import pandas as pd html_tables = pd.read_html(page) ``` Let's cover both examples in more detail as context is important. Nice video on the topic: [Easily extract tables from websites with pandas and python](https://www.youtube.com/watch?v=OXA%5FZD1gR6A&ref=datascientyst.com) Notebook: [Scrape wiki tables with pandas and python.ipynb](https://github.com/softhints/python/blob/master/notebooks/Scrape%20wiki%20tables%20with%20pandas%20and%20python.ipynb?ref=datascientyst.com) ## 1: Extract tables from PDF with Python In this example we will **extract multiple tables from remote PDF file**: [china.pdf](https://github.com/tabulapdf/tabula-java/blob/master/src/test/resources/technology/tabula/china.pdf?ref=datascientyst.com). We will use library called: [tabula-py](https://pypi.org/project/tabula-py/?ref=datascientyst.com) which can be installed by: ```bash pip install tabula-py ``` The .pdf file contains 2 table: - smaller one - bigger one with merged cells ```python from tabula import read_pdf file = 'https://raw.githubusercontent.com/tabulapdf/tabula-java/master/src/test/resources/technology/tabula/china.pdf' df_temp = read_pdf(file, stream=True) ``` After reading the data we can get a list of DataFrames which contain table data. Let's check the first one: | | FLA Audit Profile | Unnamed: 0 | | - | -------------------- | ------------------------------------------------- | | 0 | Country | China | | 1 | Factory name | 01001523B | | 2 | IEM | BVCPS (HK), Shen Zhen Office | | 3 | Date of audit | May 20-22, 2003 | | 4 | PC(s) | adidas-Salomon | | 5 | Number of workers | 243 | | 6 | Product(s) | Scarf, cap, gloves, beanies and headbands | | 7 | Production processes | Sewing, cutting, packing, embroidery, die-cutting | Which is the exact match of the first table from the PDF file. ![read-pdf-table-python-tabula](https://datascientyst.com/content/images/2022/09/read-pdf-table-python-tabula.png) While the second one is a bit weird. The reason is because of the merged cells which are extracted as `NaN` values: | | Unnamed: 0 | Unnamed: 1 | Unnamed: 2 | Findings | Unnamed: 3 | | - | -------------------------- | ----------------------------- | ------------- | ------------------ | ---------- | | 0 | FLA Code/ Compliance issue | Legal Reference / Country Law | FLA Benchmark | Monitor's Findings | NaN | | 1 | 1\. Code Awareness | NaN | NaN | NaN | NaN | | 2 | 2\. Forced Labor | NaN | NaN | NaN | NaN | | 3 | 3\. Child Labor | NaN | NaN | NaN | NaN | | 4 | 4\. Harassment or Abuse | NaN | NaN | NaN | NaN | ![read-pdf-table-python-tabula-merged-cells](https://datascientyst.com/content/images/2022/09/read-pdf-table-python-tabula-merged-cells.png) How to workaround this problem we will see in the next step. Some cells are extracted to multiple rows as we can see from the image: ## 2: Extract tables from PDF - keep format Often tables in PDF files have: - strange format - merged cells - strange symbols Most libraries and software are not able to extract them in a reliable way. To **extract complex table from PDF files with Python and Pandas** we will do: - download the file (it's possible without download) - convert the PDF file to HTML - extract the tables with Pandas ### 2.1 Convert PDF to HTML First we will download the file from: [china.pdf](https://github.com/tabulapdf/tabula-java/blob/master/src/test/resources/technology/tabula/china.pdf?ref=datascientyst.com). Then we will convert it to HTML with the library: [pdftotree](https://pypi.org/project/pdftotree/?ref=datascientyst.com). ```python import pdftotree page = pdftotree.parse('china.pdf', html_path=None, model_type=None, model_path=None, visualize=False) ``` library can be installed by: ```bash pip install pdftotree ``` ### 2.2 Extract tables with Pandas Finally we can read all the tables from this page with Pandas: ```python import pandas as pd html_tables = pd.read_html(page) html_tables[1] ``` Which will give us better results in comparison to `tabula-py` ![read-pdf-table-python-pandas-merged-cells](https://datascientyst.com/content/images/2022/09/read-pdf-table-python-pandas-merged-cells.png) ### 2.3 HTMLTableParser As alternatively to Pandas, we can use the library: [html-table-parser-python3](https://pypi.org/project/html-table-parser-python3/?ref=datascientyst.com) to parse the HTML tables to Python lists. ```python from html_table_parser.parser import HTMLTableParser p = HTMLTableParser() p.feed(page) print(p.tables[0]) ``` it convert the HTML table to Python list: ``` [['', ''], ['Country', 'China'], ['Factory name', '01001523B'], ['IEM', 'BVCPS (HK), Shen Zhen Office'], ['Date of audit', 'May 20-22, 2003'], ['PC(s)', 'adidas-Salomon'], ['Number of workers', '243'], ['Product(s)', 'Scarf, cap, gloves, beanies and headbands']] ``` Now we can convert the list to Pandas DataFrame: ```python import pandas as pd pd.DataFrame(p.tables[1]) ``` To install this library we can do: ```bash pip install html-table-parser-python3 ``` There are two differences to Pandas: - returns list of values - instead of NaN values - there are empty strings ## 3\. Python Libraries for extraction from PDF files Finally let's find a **list of useful Python libraries which can help in PDF parsing and extraction**: ### 3.1 Python PDF parsing - [tabula-py](https://pypi.org/project/tabula-py/?ref=datascientyst.com) \- Simple wrapper for tabula-java, read tables from PDF into DataFrame - [tabula-py example notebook](https://nbviewer.org/github/chezou/tabula-py/blob/master/examples/tabula%5Fexample.ipynb?ref=datascientyst.com) - [camelot-py](https://pypi.org/project/camelot-py/?ref=datascientyst.com) \- PDF Table Extraction for Humans - [pdfminer](https://pypi.org/project/pdfminer/?ref=datascientyst.com) \- PDF parser and analyzer - [PyPDF2](https://pypi.org/project/PyPDF2/?ref=datascientyst.com) \- A pure-python PDF library capable of splitting, merging, cropping, and transforming PDF files ### 3.2 Parse HTML tables - [html-table-parser-python3](https://pypi.org/project/html-table-parser-python3/?ref=datascientyst.com) \- parse HTML tables with Python 3 to list of values - [tablextract](https://pypi.org/project/tablextract/?ref=datascientyst.com) \- extracts the information represented in any HTML table - [pdftotree](https://pypi.org/project/pdftotree/?ref=datascientyst.com) \- convert PDF into hOCR with text, tables, and figures being recognized and preserved. - [pandas.read\_html](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fhtml.html?ref=datascientyst.com) - [html-table-extractor](https://pypi.org/project/html-table-extractor/?ref=datascientyst.com) \- A python library for extracting data from html table - [py-html-table](https://pypi.org/project/py-html-table/?ref=datascientyst.com) \- Python library to extract data from HTML Tables with rowspan ### 3.3 Example PDF files Finally you can find example PDF files where you can test table extraction with Python and Pandas: [tabula test PDF files](https://github.com/tabulapdf/tabula-java/tree/master/src/test/resources/technology/tabula?ref=datascientyst.com) ![](https://datascientyst.com/content/images/2022/09/extract-table-from-pdf-with-python-pandas.png) ### How to Add Border to Pandas DataFrame ( HTML Table) URL: https://datascientyst.com/add-border-pandas-dataframe-html-table/ Last updated: 2022-09-29T05:48:04.000Z This tutorial explains how to add borders to Pandas DataFrame. DataFrame will be rendered as an HTML table with borders. Adding borders to Pandas DataFrame is very useful when we work with the multi-index ## Setup Suppose we have the next DataFrame: ```python import pandas as pd df = pd.DataFrame( {"Grade": ["A", "B", "A", "C"]}, index=[ ["11", "11", "12", "12"], ["21", "22", "21", "22"], ["31", "32", "33", "34"] ] ) ``` which looks like: | | | | Grade | | -- | -- | -- | ----- | | 11 | 21 | 31 | A | | 22 | 32 | B | | | 12 | 21 | 33 | A | | 22 | 34 | C | | ## Step 1: Add inner borders to DataFrame To add inner borders in Pandas we can use the method `.style.set_table_styles()`. ```python df.style.set_table_styles([{'selector': 'td', 'props': [('font-size', '12pt'),('border-style','solid'),('border-width','1px')]}]) ``` We need to specify 2 values: - `selector` - `tr` \- table rows - `td` \- table cells - `th` \- table headers - `props` - `border-style` \- type of table border - `border-width` \- border size As we can see CSS can be applied to Pandas DataFrame. So anything valid as a CSS property can be passed to a DataFrame. This converts DataFrame to a good looking HTML table with borders. | | | | Grade | time | | -- | -- | -- | ----- | ---- | | 11 | 21 | 31 | A | 4 | | 22 | 32 | B | 5 | | | 12 | 21 | 33 | A | 6 | | 22 | 34 | C | 3 | | ## Step 2: Add outer borders to DataFrame Outer borders can be added we can be added by same method: ```python df.style.set_table_styles([{'selector': 'th', 'props': [('font-size', '12pt'),('border-style','solid'),('border-width','1px')]}]) ``` In this case we use the selector - 'th' - in order to add outer borders. | | | | Grade | time | | -- | -- | -- | ----- | ---- | | 11 | 21 | 31 | A | 4 | | 22 | 32 | B | 5 | | | 12 | 21 | 33 | A | 6 | | 22 | 34 | C | 3 | | ## Step 3: Multiple border selectors We can combine multiple selectors by using comma - `'selector': 'th,td',`. So to add borders to the whole DataFrame we can use: ```python df.style.set_table_styles([{'selector': 'th,td', 'props': [('font-size', '12pt'),('border-style','solid'),('border-width','1px')]}]) ``` | | | | Grade | time | | -- | -- | -- | ----- | ---- | | 11 | 21 | 31 | A | 4 | | 22 | 32 | B | 5 | | | 12 | 21 | 33 | A | 6 | | 22 | 34 | C | 3 | | ![](https://datascientyst.com/content/images/2022/09/add-border-pandas-dataframe-html-table.png) ## Step 4: Table colors and more styles Adding color is possible by property: `background-color`. For example: ```python styler0 = df.style styler0.set_table_attributes('style="font-size: 25px"') styler0.set_table_styles([{'selector': '*', 'props': [('color', 'black'),('border-style','solid'),('border-width','1px')]}, {'selector': 'th', 'props': [('background-color', color1)]}]) ``` First we copy the `.style` object. Then we set the font size for the whole DataFrame. Next we modify all properties of the table and finally the table borders. ## Step 5: DataFrame to HTML code To get the code from the HTML you can check this article: [Render Pandas DataFrame As HTML Table Keeping Style](https://datascientyst.com/render-pandas-dataframe-html-table-keeping-style/) In short we can use for the DataFrame without styles: ```python df.to_html() ``` or render DataFrame as HTML code with styles: ```python display_html(df._repr_html_()) ``` ### How to Use .loc and Multi-Index in Pandas URL: https://datascientyst.com/use-loc-and-multi-index-in-pandas/ Last updated: 2023-03-17T23:07:19.000Z In this tutorial, we'll see how to **select values with `.loc()` on multi-index in Pandas DataFrame.** Here are quick solutions for selection on multi-index: **(1) Select first level of MultiIndex** ```python df2.loc['11', :] ``` **(2) Select columns - MultiIndex** ```python df.loc[0, ('company A', ['rank'])] ``` **(3) Conditional selection on level of MultiIndex** ```python mask = (df2.index.get_level_values(0)=='11') | (df2.index.get_level_values(1)=='22') df2[mask] ``` ## Setup For this article we're going to use two DataFrames: - first one with multi index on columns - second one with multi index on rows ```python import pandas as pd cols = pd.MultiIndex.from_tuples([('company A', 'rank'), ('company A', 'points'), ('company B', 'rank'), ('company B', 'points')]) df = pd.DataFrame([[1,2,3,4], [2,3, 3,4]], columns=cols) ``` data: | | company A | company B | | | | - | --------- | --------- | ---- | ------ | | | rank | points | rank | points | | 0 | 1 | 2 | 3 | 4 | | 1 | 2 | 3 | 3 | 4 | The second one: ```python df2 = pd.DataFrame( {"Grade": ["A", "B", "A", "C"]}, index=[ ["11", "11", "12", "12"], ["21", "22", "21", "22"], ["31", "32", "33", "34"] ] ) ``` data | | | | Grade | | -- | -- | -- | ----- | | 11 | 21 | 31 | A | | 22 | 32 | B | | | 12 | 21 | 33 | A | | 22 | 34 | C | | ## Step 1: .loc() and MultiIndex Pandas method `.loc()` can select on multi-index. To find out what are the index values we can use method: `df2.index` which will give us: ``` MultiIndex([('11', '21', '31'), ('11', '22', '32'), ('12', '21', '33'), ('12', '22', '34')], ) ``` To select values from the multi index above we can use following syntax: ```python df2.loc[('11', '21', '31'), :] ``` which give us: ``` Grade A Name: (11, 21, 31), dtype: object ``` ## Step 2: Select first level of multi-index To select first level of multiindex we can use method `loc()` and provide list of values: ```python df2.loc[['11']] ``` result: | | | | Grade | | -- | -- | -- | ----- | | 11 | 21 | 31 | A | | 22 | 32 | B | | To select multiple values from the first level we can use: ```python df2.loc[['11', '12']] ``` ## Step 3: Select second level of multi-index To select second or N-th level from multi index in Pandas DataFrame we can use `slice(None)` or method `get_level_values()`: ```python df2[df2.index.get_level_values(1)=='21'] ``` or ```python sel = (slice(None), ['21'], slice(None)) df2.loc[sel] ``` result in selection of the second level of this multi-index: | | | | Grade | | -- | -- | -- | ----- | | 11 | 21 | 31 | A | Alternatively we can create IndexSlice object: ```python idx = pd.IndexSlice df2.loc[idx[:,['21'],:],:] ``` to get the same result. ## Step 4: Conditional selection on multi-index For conditional selection on multi index in pandas we can use method `get_level_values()` and mask: ```python mask = (df2.index.get_level_values(0)=='11') | (df2.index.get_level_values(1)=='22') df2[mask] ``` In this way we can combine multiple conditions: - select first level - value '11' - or second level - value '22' | | | | Grade | | -- | -- | -- | ----- | | 11 | 21 | 31 | A | | 22 | 32 | B | | | 12 | 22 | 34 | C | ![](https://datascientyst.com/content/images/2022/09/use-loc-and-multi-index-in-pandas-datascientyst.png) ## Step 5: Query selection on multi-index If our multi-index has named levels we can use queries to select data: ```python df.query('level1 == "11" | level2 == "21"') ``` ## Step 6: Select multi-index column All from above apply to columns. For columns we need to pass criteria as second value: ```python df.loc[0, ('company A', ['rank'])] ``` result: ``` company A rank 1 Name: 0, dtype: int64 ``` ## Conclusion In this article, we looked at different solutions for selection and quering data from Pandas Multi-Index. We focused on row/index selection, column selection is exactly the same. **We covered conditional selection, selection from first or second level and queries on Multi-Index DataFrame.** ### 433-project URL: https://datascientyst.com/project/ Last updated: 2022-09-22T14:23:48.000Z Data Science Projects ### How to Count Values in Pandas DataFrame URL: https://datascientyst.com/count-values-pandas-dataframe/ Last updated: 2022-09-22T14:10:15.000Z In this tutorial, we're going to **count values in Pandas DataFrame**. Check this article for most common values in DataFrame: [Get most frequent values in Pandas DataFrame](https://datascientyst.com/get-most-frequent-values-pandas-dataframe/) ## Setup Suppose we have a dataframe with data: ```python import pandas as pd data = [('A',1, 0, 3, 1), ('A',1, 2, 5, 1), ('B',2, 1, 4, 3), ('B',3, 1, 0, 3), ('C',4, 3, 1, 2)] cols = ('col_1', 'col_2', 'col_3', 'col_4', 'col_5' ) df = pd.DataFrame(data, columns = cols) ``` data: | | col\_1 | col\_2 | col\_3 | col\_4 | col\_5 | | - | ------ | ------ | ------ | ------ | ------ | | 0 | A | 1 | 0 | 3 | 1 | | 1 | A | 1 | 2 | 5 | 1 | | 2 | B | 2 | 1 | 4 | 3 | | 3 | B | 3 | 1 | 0 | 3 | | 4 | C | 4 | 3 | 1 | 2 | ## Step 1: Count values in Pandas Column To **count values in single Pandas column** we can use method `value_counts()`: ```python df['col_1'].value_counts() ``` The result is count of most frequent values sorted in ascending order: ``` A 2 B 2 C 1 Name: col_1, dtype: int64 ``` ## Step 2: Count values in Multiple columns **Count values in multiple Pandas columns** can be down with method `.value_counts()`: ```python df[['col_1', 'col_2']].value_counts() ``` The result is is most common values per group/columns: ``` col_1 col_2 A 1 2 B 2 1 3 1 C 4 1 dtype: int64 ``` ## Step 3: Count values in Pandas DataFrame To find the **most common values in whole DataFrame** we can combine: - `melt` - `value_counts` ```python df.melt()['value'].value_counts() ``` So we get the\*\* count of all values for this DataFrame\*\*: ``` 1 7 3 5 2 3 A 2 B 2 4 2 0 2 C 1 5 1 Name: value, dtype: int64 ``` ![](https://datascientyst.com/content/images/2022/09/count-values-pandas-dataframe-column.png) ## Conclusion In this short article we covered **how to count values in Pandas DataFrame**. We saw how to count values in single or multiple columns. Finally with covered how to get most common values for the whole DataFrame ### How to Extract Year-Week from DateTime in Pandas URL: https://datascientyst.com/extract-year-week-datetime-pandas/ Last updated: 2022-09-02T21:56:08.000Z In this tutorial, we'll see how to **extract year-week from datetime column in Pandas DataFrame.** So at the end we will get: - `2019-10-14 00:00:00 -> 2019-w42` - `2018-10-15 -> 2018-w42` - `2018-10-15 -> 2018-42` **(1) Extract year-week as date - strftime** ```python df['col'].dt.strftime('%Y-w%V') ``` **(2) Extract year-week as string** ```python df['y'].apply(str) + '-' + df['w'].apply(str) ``` You can find more about Pandas dates extraction - [How to Convert DateTime to Day of Week(name and number) in Pandas](https://datascientyst.com/convert-datetime-day-of-week-name-number-in-pandas/) - [How to Extract Month and Year from DateTime column in Pandas](https://datascientyst.com/extract-month-and-year-datetime-column-in-pandas/) ## Setup For this example we will use DataFrame like: ```python import pandas as pd data = { 'age': [25, 30, 40, 35, 20, 40, 22], 'start_date': ["2019-10-14", "2018-10-15","2020-7-15", "2020-10-6","2020-03-8","2015-10-14","2011-12-18"], 'person': ['Tim', 'Jim', 'Kim', 'Bim', 'Dim', 'Sim', 'Lim'] } df = pd.DataFrame(data) ``` In this DataFrame there is a date column: | | age | start\_date | person | | - | --- | ----------- | ------ | | 0 | 25 | 2019-10-14 | Tim | | 1 | 30 | 2018-10-15 | Jim | | 2 | 40 | 2020-07-15 | Kim | | 3 | 35 | 2020-10-06 | Bim | | 4 | 20 | 2020-03-08 | Dim | | 5 | 40 | 2015-10-14 | Sim | | 6 | 22 | 2011-12-18 | Lim | If data is read as a string we need to convert it to datetime by: ```python df['start_date'] = pd.to_datetime(df['start_date']) ``` In order to avoid errors like: ``` AttributeError: Can only use .dt accessor with datetimelike values ``` ## 1: Extract year-week in Pandas Let's start by extracting the year-week column from datetime in Pandas. The most convenient and shortest way is by using `strftime`: ```python df['start_date'].dt.strftime('%Y-w%V') ``` which will result into: ``` 0 2019-w42 1 2018-w42 2 2020-w29 3 2020-w41 4 2020-w10 5 2015-w42 6 2011-w50 Name: start_date, dtype: object ``` ![](https://datascientyst.com/content/images/2022/09/extract-year-week-datetime-pandas-datascientyst.png) ## 2: Extract year-week as string Alternatively if we have columns for year and week we can combine them by `+` operator. ```python df['w'] = df['start_date'].dt.isocalendar().week df['y'] = df['start_date'].dt.isocalendar().year df['yw'] = df['y'].apply(str) + '-' + df['w'].apply(str) ``` The output is year-week pairs: ``` 0 2019-42 1 2018-42 2 2020-29 3 2020-41 4 2020-10 5 2015-42 6 2011-50 Name: yw, dtype: object ``` *Rule of thumb:* Be sure that columns are converted to string with `.apply(str)` \- to avoid errors like: ``` TypeError: can only perform ops with numeric values ``` **Pro Tip 1** Storing year-week values as a string might impact data analysis. For example, sorting and visualization will not work as expected. ## 3: Week number formats & directives There are several possible ways to extract week numbers: - `%V` \- 1..53 - starting Monday (ISO standard) - `%W` \- 0..53 - start on Sunday - `%U` \- 0..53 - Monday first day of week More info can be found on: - [time.strftime()](https://docs.python.org/3/library/time.html?ref=datascientyst.com#time.strftime) - [pandas.Period.strftime](https://pandas.pydata.org/docs/reference/api/pandas.Period.strftime.html?ref=datascientyst.com) ## Conclusion In this article, we learned how to week and year-week in Pandas. We started by date format extraction, then we covered string extraction with potential problems. Finally we covered different week formats in Python and `strftime`. ### How to Melt Pandas DataFrame - pd.melt in Examples URL: https://datascientyst.com/use-melt-pandas-dataframe-pd-melt-examples/ Last updated: 2022-09-01T21:17:32.000Z In this quick tutorial, we'll see how to use melt in Pandas. We'll first look into basic pd.melt usage, then `pd.melt()` parameters, and finally some advanced examples and alternatives of melt in Pandas and Python. In short we can do: **(1) pd.melt() in Pandas** ```python pd.melt(df, id_vars=['A'], value_vars=['B']) ``` **(2) pd.melt() and MultiIndex** ```python pd.melt(df, id_vars=[('A', 'D')], value_vars=[('B', 'E')]) ``` ## Setup In this article we will use DataFrame which has information for population of countries: ```python import pandas as pd data = [{"Country":"China","1950":"562,580","1955":"607,047","1960":"651,340","1980":"987,822","continent":"Asia"}, {"Country":"India","1950":"369,881","1955":"404,268","1960":"445,394","1980":"684,888","continent":"Asia"}, {"Country":"United States","1950":"151,869","1955":"165,070","1960":"179,980","1980":"227,225","continent":"N. America"}, {"Country":"Indonesia","1950":"82,979","1955":"90,255","1960":"100,146","1980":"150,322","continent":"Asia"}, {"Country":"Russia","1950":"101,937","1955":"111,126","1960":"119,632","1980":"139,039","continent":"Europe"}, {"Country":"Brazil","1950":"53,444","1955":"61,652","1960":"71,412","1980":"121,064","continent":"S. America"}] pd.DataFrame(data) ``` Final data format is perfect example of where Pandas melt function is useful: | | Country | 1950 | 1955 | 1960 | 1980 | continent | | - | ------------- | ------- | ------- | ------- | ------- | ---------- | | 0 | China | 562,580 | 607,047 | 651,340 | 987,822 | Asia | | 1 | India | 369,881 | 404,268 | 445,394 | 684,888 | Asia | | 2 | United States | 151,869 | 165,070 | 179,980 | 227,225 | N. America | | 3 | Indonesia | 82,979 | 90,255 | 100,146 | 150,322 | Asia | | 4 | Russia | 101,937 | 111,126 | 119,632 | 139,039 | Europe | | 5 | Brazil | 53,444 | 61,652 | 71,412 | 121,064 | S. America | ## 1: What is melt in Pandas The picture below shows melt function in action ![](https://datascientyst.com/content/images/2022/09/use-melt-pandas-dataframe-pd-melt-examples-datascientyst.png) There are 2 important parameters of this method: - `id_vars` \- identifier variables - `value_vars` \- measured variables, which are "melt" or "unpivoted" to row axis (non-identifier columns) - `value` \- is the column values - `variable` \- the column names So the melt function will turn multiple columns - `value_vars` \- to rows. There will be two non-identifier columns. There are `id_vars` \- identifier variables which will be considered as identifier columns. **Info** In short pd.melt() unpivots data. It makes it easier to filter, compare, visualize and use in DB like style. ## 2: Why is melt useful Pandas melt function is useful when data is in form of: - pivot - pivot table - cross table And we need to convert the multiple columns to rows. Looking at the data above we might want to find the top 5 most populated pairs - year and country. Using `melt()` will help us to get this information much faster. The reverse operation of melt can be found on: [Opposite of Melt in Python and Pandas](https://datascientyst.com/opposite-of-melt-python-pandas/) ## 3: Pandas melt example Let's see how to use the melt function. First we will identify the parameters: - `id_vars=['Country', 'continent']` - `value_vars=['1950', '1955']` ```python pd.melt(df, id_vars=['Country', 'continent'], value_vars=['1950', '1955']) ``` This would result into: | | Country | continent | variable | value | | -- | ------------- | ---------- | -------- | ------- | | 0 | China | Asia | 1950 | 562,580 | | 1 | India | Asia | 1950 | 369,881 | | 2 | United States | N. America | 1950 | 151,869 | | 3 | Indonesia | Asia | 1950 | 82,979 | | 4 | Russia | Europe | 1950 | 101,937 | | 5 | Brazil | S. America | 1950 | 53,444 | | 6 | China | Asia | 1955 | 607,047 | | 7 | India | Asia | 1955 | 404,268 | | 8 | United States | N. America | 1955 | 165,070 | | 9 | Indonesia | Asia | 1955 | 90,255 | | 10 | Russia | Europe | 1955 | 111,126 | | 11 | Brazil | S. America | 1955 | 61,652 | So the output has 4 columns: - 2 identifier columns - Country - continent - 2 non identifier columns - variable - value If we add more `value_vars` columns the number of the columns will be the same. Only `id_vars` change the number of the output columns. ## 4: Pandas melt - change names If we like to change the non-identifier columns we can use 2 parameters: - `var_name` - `value_name` and set the new names for the result columns: ```python pd.melt(df, id_vars=['Country', 'continent'], value_vars=['1950', '1955'], var_name ='year', value_name ='value') ``` ## 5: Pandas melt - parameters There are several parameters of melt function. Most of them were mentioned in the article. Two remains to be described: - `col_level` \- specifies the level to melt (MultiIndex) - `ignore_index` \- ignore the original index ## 6: Pandas melt - MultiIndex Melt function can be used for MultiIndex DataFrame. To simulate melt function for MultiIndex we will turn the columns into MultiIndex by: ```python df = pd.concat({'year': df}, names=['Firstlevel'], axis=1) ``` More about: [How to Add a Level to Index in Pandas DataFrame](https://datascientyst.com/add-level-index-pandas-dataframe/) We can check the columns by: ```python df.columns ``` result: ``` MultiIndex([('year', 'Country'), ('year', '1950'), ('year', '1955'), ('year', '1960'), ('year', '1980'), ('year', 'continent')], names=['Firstlevel', None]) ``` Melt can be invoked on MultiIndex columns by: ```python pd.melt(df, id_vars=[('year', 'Country')], value_vars=[('year', '1950'), ('year', '1955')]) ``` The result is: | | (year, Country) | Firstlevel | None | value | | -- | --------------- | ---------- | ---- | ------- | | 0 | China | year | 1950 | 562,580 | | 1 | India | year | 1950 | 369,881 | | 2 | United States | year | 1950 | 151,869 | | 3 | Indonesia | year | 1950 | 82,979 | | 4 | Russia | year | 1950 | 101,937 | | 5 | Brazil | year | 1950 | 53,444 | | 6 | China | year | 1955 | 607,047 | | 7 | India | year | 1955 | 404,268 | | 8 | United States | year | 1955 | 165,070 | | 9 | Indonesia | year | 1955 | 90,255 | | 10 | Russia | year | 1955 | 111,126 | | 11 | Brazil | year | 1955 | 61,652 | ## Conclusion Pandas melt is a very useful function for reshaping DataFrames. It can multiple columns to rows. We saw how it works, why it's important and several examples. The function can help "unpivoting" data. ### How to Append Pandas Series to DataFrame URL: https://datascientyst.com/append-pandas-series-dataframe/ Last updated: 2024-01-19T23:06:55.000Z To append Series to DataFrame in Pandas we have several options: **(1) Append Series to DataFrame by `pd.concat`** ```python pd.concat([df, humidity.to_frame()], axis=1) ``` **(2) Append Series with - `append` (will be deprecated)** ```python df.append(humidity, ignore_index=True) ``` **(3) Append the Series to DataFrame with assign** ```python df['humidity'] = humidity ``` P.S. Not that since Pandas 2.0 method append is deprecated which results into error: "'DataFrame' object has no attribute 'append'". To find how to fix the error: [Why is append not working in pandas?](https://datascientyst.com/fix-attributeerror-dataframe-object-has-no-attribute-append-pandas/) ## Setup In the examples, we'll use the following setup, which consists of 1 DataFrame and 2 Series: ```python import pandas as pd import matplotlib.pyplot as plt data={'day': [1, 2, 3, 4, 5], 'temp': [9, 8, 6, 13, 10]} df = pd.DataFrame(data) ``` data is: | | day | temp | | - | --- | ---- | | 0 | 1 | 9 | | 1 | 2 | 8 | | 2 | 3 | 6 | | 3 | 4 | 13 | | 4 | 5 | 10 | The series are: ```python ser_col = pd.Series(data=[0.89, 0.86, 0.54, 0.73, 0.45], name='humidity') ``` which contains data for a new column: ``` 0 0.89 1 0.86 2 0.54 3 0.73 4 0.45 Name: humidity, dtype: float64 ``` And a Series which contains data for a new row: ```python ser_row = pd.Series({'day': 6, 'temp': 15}) ``` result: ``` day 6 temp 15 dtype: int64 ``` ## 1: Append Series to DataFrame - pd.concat as column Let's start by appending Pandas Series to DataFrame by method `pd.concat()`. We will add the Series to the DataFrame as a new column - this is possible by using parameter - `axis=1`: ```python pd.concat([df, ser_col.to_frame()], axis=1) ``` This will append the Series as a new column to the DataFrame: | | day | temp | humidity | | - | --- | ---- | -------- | | 0 | 1 | 9 | 0.89 | | 1 | 2 | 8 | 0.86 | | 2 | 3 | 6 | 0.54 | | 3 | 4 | 13 | 0.73 | | 4 | 5 | 10 | 0.45 | More information for this method: [pandas.concat](https://pandas.pydata.org/docs/reference/api/pandas.concat.html?ref=datascientyst.com) ## 2: Append Series to DataFrame - pd.concat as row Next, we'll add the Series as a new row to a DataFrame. The method concat can be used once again: ```python pd.concat([df, ser_row.to_frame().T], ignore_index=True) ``` We can ignore the index by using - `ignore_index=True`: | | day | temp | | - | --- | ---- | | 0 | 1 | 9 | | 1 | 2 | 8 | | 2 | 3 | 6 | | 3 | 4 | 13 | | 4 | 5 | 10 | | 5 | 6 | 15 | This is equivalent to `axis=0` ![append-pandas-series-dataframe](https://datascientyst.com/content/images/2022/10/append-pandas-series-dataframe.webp) ## 3\. Add Series to DataFrame with append Let's see how to use the method `pd.append()` which will be deprecated in future. But it's still working and can be useful in some cases. **Pro Tip 1** Method **pd.append** is deprecated since version 1.4.0: Use concat() instead. ### Append Series as a row To append Pandas Series as a new row to a DataFrame we can do: ```python df.append(ser_row,ignore_index=True) ``` The new DataFrame is: | | day | temp | | - | --- | ---- | | 0 | 1 | 9 | | 1 | 2 | 8 | | 2 | 3 | 6 | | 3 | 4 | 13 | | 4 | 5 | 10 | | 5 | 6 | 15 | More information for this method: [pandas.DataFrame.append](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.append.html?ref=datascientyst.com) ## 4\. Add Series to DataFrame as column Finally, let's take a look at a solution with assigning a new Series as a column in Pandas DataFrame. We can simply assign the Series to the DataFrame: ```python df['humidity'] = ser_col ``` The result is the a new column: | | day | temp | humidity | | - | --- | ---- | -------- | | 0 | 1 | 9 | 0.89 | | 1 | 2 | 8 | 0.86 | | 2 | 3 | 6 | 0.54 | | 3 | 4 | 13 | 0.73 | | 4 | 5 | 10 | 0.45 | **Pro Tip 1** Pandas Series can be considered as a DataFrame column - that's why we can add them to a DataFrame. ## Conclusion In this article, we looked at different solutions for appending/adding a Series to DataFrame in Pandas. We focused on the most popular solutions with `pd.concat` but also covered a few alternatives. We saw how to add a Series as a column or a row. ### How to Read JSON Files in Pandas URL: https://datascientyst.com/read-json-files-pandas/ Last updated: 2022-08-30T21:34:32.000Z In this tutorial, we'll focus on reading JSON files with Pandas and Python. We will cover reading JSON files and JSON lines( read the file as JSON object per line). **(1) Reading JSON file in Pandas** ```python pd.read_json('file.json') ``` **(2) Reading JSON line file in Pandas** ```python pd.read_json('file.json', lines=True) ``` ## Setup In our example we'll be using a DataFrame with the next data. ```python import pandas as pd data = {'day': [1, 2, 3, 4, 5, 6, 7, 8], 'temp': [9, 8, 6, 13, 10, 15, 9, 10], 'humidity': [0.89, 0.86, 0.54, 0.73, 0.45, 0.63, 0.95, 0.67]} df = pd.DataFrame(data=data) ``` We will use also a file called 'file.json' which can be exported from this DataFrame by: ```python df.to_json(orient='columns') ``` then we can read it again to DataFrame with `read_json()`: ```python pd.read_json(df.to_json(orient='columns'), orient='columns') ``` The content of 'file.json': ``` { "day":{ "0":1, "1":2, "2":3, "3":4 }, "temp":{ "0":9, "1":8, "2":6, "3":13 }, "humidity":{ "0":0.89, "1":0.86, "2":0.54, "3":0.73 } } ``` ## 1: Read JSON file with Pandas To read a JSON file named 'file.json' we can use the method `read_json()`. The official documentation is placed on link: [pandas.read\_json](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fjson.html?ref=datascientyst.com) By default the method is reading `orient='columns'`: ```python pd.read_json('file.json') ``` which is equivalent to: ```python pd.read_json('file.json', orient='columns') ``` The method can use buffer or relative path to the JSON file: ```python pd.read_json(r'../data/file.json') ``` ## 2: Read JSON lines with Pandas Pandas can read JSON lines file which is a file with JSON objects stored on separate lines: ``` {"day":1,"temp":9,"humidity":0.89} {"day":2,"temp":8,"humidity":0.86} {"day":3,"temp":6,"humidity":0.54} ``` we can use parameter - `lines=True`: ```python pd.read_json('file_lines.jl', lines=True) ``` To simulate exporting DataFrame to JSON lines and importing it back to DataFrame we need to use - `orient='records'` for the export: ```python pd.read_json(df.to_json(orient='records', lines=True), lines=True) ``` ![](https://datascientyst.com/content/images/2022/08/datascientyst-read-json-files-pandas.png) ## 3: Pandas read\_json() parameters There multiple important parameters of method `ead_json()`: - `orient` \- expected JSON string format - check next section for more info - `dtype` \- if True, infer dtypes; if a dict of column to dtype, then use those - `convert_dates` \- convert date-like columns(depends on `keep_default_dates`) - `lines` \- read JSON lines - `nrows` \- number of lines to be read ## 4: JSON formats - Pandas There are multiple options for parameter `orient` of Pandas method - read\_json: - `split` \- dict - `{‘index’ -> [index], ‘columns’ -> [columns], ‘data’ -> [values]}` - `records` \- list - `[{column -> value}, … , {column -> value}]` - `index` \- dict - `{column -> {index -> value}}` - `columns` \- dict (default for DataFrame) - `{column -> {index -> value}}` - `values` - just the values array - `table` \- dict - `{‘schema’: {schema}, ‘data’: {data}}` For more information on JSON formats and extraction you can check: [How to Export DataFrame to JSON with Pandas](https://datascientyst.com/export-dataframe-to-json-pandas/). ## 5: Read Python dict with Pandas Next let's cover two related topics: - reading Python dict with Pandas - what is the difference between Python dict and JSON So JSON vs Python dict: - Python dict - data structure (memory object) - JSON - universal (language independent) data format (string-based storage) Reading Python dict with Pandas like: ```python pd.DataFrame(d) ``` Might raise an error like: ``` ValueError: If using all scalar values, you must pass an index ``` The error can be solved by getting the items from the dictionary by `.items()`: ```python pd.DataFrame(d.items()) ``` ## 6: Read semi-structured JSON Finally we can see how to convert semi-structured JSON data into a flat table. This can be done by method - `json_normalize()`: ```python data = [ {"id": 1, "name": {"first": "mr", "last": "X"}}, {"name": {"given": "mr", "family": "Y"}} ] pd.json_normalize(data) ``` which results into Pandas DataFrame: | | id | name.first | name.last | name.given | name.family | | - | --- | ---------- | --------- | ---------- | ----------- | | 0 | 1.0 | mr | X | NaN | NaN | | 1 | NaN | NaN | NaN | mr | Y | This method can be combined with `json.load()` in order to read strange JSON formats: ```python import json df = pd.json_normalize(json.load(open("file.json", "rb"))) ``` ## 7: Read JSON files with json.load() In some cases we can use the method `json.load()` to read JSON files with Python. Then we can pass the read JSON data to Pandas DataFrame constructor like: ```python import json with open('file.json') as f: data = json.load(f) pd.DataFrame(data) ``` This option is useful for performance sake or when there are errors. ## 8\. ValueError: Invalid file path or buffer object type: The Pandas errors like ``` "ValueError: Invalid file path or buffer object type: " ``` or ``` ValueError: Invalid file path or buffer object type: ``` is raised when the method is abused like: ```python pd.read_json({'a':1}) ``` ## Conclusion In this article, we saw how to read JSON files, JSON lines objects and multiple JSON formats. We discussed alternative ways to read JSON files and how to deal with semi-structured JSON like data. ### How to Add a Level to Index in Pandas DataFrame URL: https://datascientyst.com/add-level-index-pandas-dataframe/ Last updated: 2022-08-30T14:18:30.000Z In this tutorial, we'll explore how to **add level to MultiIndex in Pandas DataFrame.** We will learn how to add levels to rows or columns. In short we can do something like: **(1) Prepend level to Index** ```python pd.concat([df], keys=['week_1'], names=['level_name']) ``` **(2) Flexible way to add new level** ```python old_idx.insert(0, 'level_name', new_ix_level) ``` **(3) Use df.set\_index to add new level to index** ```python df.set_index('level_name', append=True, inplace=True) ``` ## Setup Let's work with the following DataFrame: ```python import pandas as pd data = {'day': [1, 2, 3, 4, 5, 6, 7, 8], 'temp': [9, 8, 6, 13, 10, 15, 9, 10], 'humidity': [0.89, 0.86, 0.54, 0.73, 0.45, 0.63, 0.95, 0.67]} df = pd.DataFrame(data=data) ``` data looks like: | | day | temp | humidity | | - | --- | ---- | -------- | | 0 | 1 | 9 | 0.89 | | 1 | 2 | 8 | 0.86 | | 2 | 3 | 6 | 0.54 | | 3 | 4 | 13 | 0.73 | | 4 | 5 | 10 | 0.45 | ## 1: Add first level to Index/MultiIndex (rows) Let's start by **adding new level to index** \- so we will create MultiIndex from index in Pandas DataFrame: ```python pd.concat([df], keys=['week_1'], names=['level_name']) ``` Before this operation we have - `df.index`: ``` RangeIndex(start=0, stop=8, step=1) ``` After this operation we get: ``` MultiIndex([('week_1', 0), ('week_1', 1), ('week_1', 2), ('week_1', 3)...], names=['level_name', None]) ``` The new DataFrame with the new MultiIndex look like: | | | day | temp | humidity | | ----------- | - | --- | ---- | -------- | | level\_name | | | | | | week\_1 | 0 | 1 | 9 | 0.89 | | 1 | 2 | 8 | 0.86 | | | 2 | 3 | 6 | 0.54 | | | 3 | 4 | 13 | 0.73 | | | 4 | 5 | 10 | 0.45 | | ### Access level of MultiIndex **Pro Tip 1** To access the data from the multiIndex we can do: df.loc[('week_1', 0)] This will result into: ``` day 1.00 temp 9.00 humidity 0.89 Name: (week_1, 0), dtype: float64 ``` ## 2: Add second level to Index/MultiIndex (rows) Alternatively **we can add second(inner) level to create MultiIndex** by: - creating new column with values - set the new column as index ```python df['level_name'] = 'week_1' df.set_index('level_name', append=True, inplace=True) ``` The result is: | | | day | temp | humidity | | - | ----------- | --- | ---- | -------- | | | level\_name | | | | | 0 | week\_1 | 1 | 9 | 0.89 | | 1 | week\_1 | 2 | 8 | 0.86 | | 2 | week\_1 | 3 | 6 | 0.54 | | 3 | week\_1 | 4 | 13 | 0.73 | | 4 | week\_1 | 5 | 10 | 0.45 | This time the DataFrame MultiIndex is: ``` MultiIndex([(0, 'week_1'), (1, 'week_1'), (2, 'week_1'), (3, 'week_1')..], names=[None, 'level_name']) ``` ## 3: Add first level to columns(index) To **add new level to the columns** and create a MultiIndex we can use the following code: ```python df = pd.concat([df], keys=['week_1'], names=['level_name'], axis=1) ``` This will change our DataFrame to: | level\_name | meteo | | | | ----------- | ----- | ---- | -------- | | | day | temp | humidity | | 0 | 1 | 9 | 0.89 | | 1 | 2 | 8 | 0.86 | | 2 | 3 | 6 | 0.54 | | 3 | 4 | 13 | 0.73 | | 4 | 5 | 10 | 0.45 | Checking the column index - `df.columns`: ``` MultiIndex([('meteo', 'day'), ('meteo', 'temp'), ('meteo', 'humidity')], names=['level_name', None]) ``` ![](https://datascientyst.com/content/images/2022/08/add-level-index-pandas-dataframe.png) ## 4: Flexible way to add new level of Index (rows/columns) A generic solution to add **new levels of Index or MultiIndex in Pandas DataFrame** is by: - converting the index to DataFrame - update the index - change back to index/MultiIndex So the code below shows all the steps: ```python new_ix_level = ['week_1'] * 4 + ['week_2'] * 4 # new values for the level old_idx = df.index.to_frame() # convert to DataFrame old_idx.insert(0, 'level_name', new_ix_level) # add new level df.index = pd.MultiIndex.from_frame(old_idx) ``` The new level of the MultiIndex has several different values: ``` MultiIndex([('week_1', 0), ('week_1', 1), ('week_1', 2), ('week_1', 3), ('week_2', 4), ('week_2', 5)..], names=['level_name', 0]) ``` and the DataFrame is: | | day | temp | humidity | | - | --- | ---- | -------- | | 0 | 1 | 9 | 0.89 | | 1 | 2 | 8 | 0.86 | | 2 | 3 | 6 | 0.54 | | 3 | 4 | 13 | 0.73 | | 4 | 5 | 10 | 0.45 | ## 5: Add multiple levels to index Finally let's check how we **can add multiple levels to index in Pandas DataFrame**. We will use the last generic solution in order to add two levels to the index - this will create MultiIndex with 3 levels: ```python new_ix_level = ['week_1'] * 4 + ['week_2'] * 4 new_ix_level_1 = ['2022'] * 8 old_idx = df.index.to_frame() old_idx.insert(0, 'week', new_ix_level) old_idx.insert(0, 'year', new_ix_level_1) df.index = pd.MultiIndex.from_frame(old_idx) ``` result: | | | | day | temp | humidity | | ------- | ------- | -- | ---- | ---- | -------- | | year | week | 0 | | | | | 2022 | week\_1 | 0 | 1 | 9 | 0.89 | | 1 | 2 | 8 | 0.86 | | | | 2 | 3 | 6 | 0.54 | | | | 3 | 4 | 13 | 0.73 | | | | week\_2 | 4 | 5 | 10 | 0.45 | | ## Conclusion In this post, we covered multiple ways to add levels to Index or MultiIndex in Pandas DataFrame. We saw how to add the first or last level and a generic solution for adding different values - in the new level. Finally we also saw how to add multiple levels at once. ### How to Merge CSV Files with Python (Pandas DataFrame) URL: https://datascientyst.com/merge-csv-files-python-pandas-dataframe/ Last updated: 2022-08-28T20:44:39.000Z In this short guide, we're going to **merge multiple CSV files into a single CSV file with Python**. We will also see how to **read multiple CSV files - by wildcard matching - to a single DataFrame**. The code to **merge several CSV files matched by pattern to a file or Pandas DataFrame is**: ```python import glob for f in glob.glob('file_*.csv'): df_temp = pd.read_csv(f) ``` ## Setup Suppose we have multiple CSV files like: - file\_202201.csv - file\_202202.csv - file\_202203.csv into single CSV file like: `merged.csv` ![](https://datascientyst.com/content/images/2022/08/merge-csv-files-python-pandas-dataframe.png) ## 1: Merge CSV files to DataFrame To merge multiple CSV files to a DataFrame we will use the Python module - `glob`. The module allow us to search for a file pattern with wildcard - `*`. ```python import pandas as pd import glob df_files = [] for f in glob.glob('file_*.csv'): df_temp = pd.read_csv(f) df_files.append(df_temp) df = pd.concat(df_files) ``` How does the code work? - All files which match the pattern will be iterated in random order - Temporary DataFrame is created for each file - The temporary DataFrame is appended to list - Finally all DataFrames are merged into a single one ## 2: Read CSV files without header To **skip the headers** for the CSV files we can use parameter: `header=None` ```python read_csv('file_*.csv', header=None) ``` To add the headers only for the first file we can: - read the first file with headers - drop duplicates (keep first) - set column names All depends on the context. ## 3: Read sorted CSV files Module `glob` reads files without order. To ensure the **correct order of the read CSV files** we can use `sorted`: ```python sorted(glob.glob('file_*.csv')) ``` This ensures that the final output CSV file or DataFrame will be loaded in a certain order. Alternatively we can use parameters: `ignore_index=True, , sort=True` for Pandas method `concat`: ```python merged_df = pd.concat(all_df, ignore_index=True, sort=True) ``` ## 4: Change CSV separator We can control what is the **separator symbol for the CSV** files by using parameter: ```python sep = '\t' ``` Default one is comma. ## 5: Keep trace of CSV files If we like to **keep trace** of each row loaded - from which CSV file is coming we can use: `df_temp['file'] = f.split('/')[-1]`: This will data a new column to each file with trace - the file name origin. ## 6: Merge CSV files to single with Python Finally we can save the result into a single CSV file from Pandas Dataframe by: ```python df_merged.to_csv("merged.csv") ``` ## 7\. Full Code Finally we can find the full example with most options mentioned earlier: ```python import pandas as pd import glob df_files = [] for f in sorted(glob.glob('file_*.txt')): df_temp = pd.read_csv(f, header=None, index_col=False, sep='\t') df_temp['file'] = f.split('/')[-1] df_files.append(df_temp) df_merged = pd.concat(df_files) df_merged.to_csv("merged.csv") ``` ## Conclusion We saw how to read multiple CSV files with Pandas and Python. Different options were covered like: - changing separator for `read_csv` - keeping trace of the source file - sorting files in certain order - skipping headers ### 131-read-json URL: https://datascientyst.com/read_dict/ Last updated: 2022-08-28T09:27:28.000Z read\_json() ### 214-to-json URL: https://datascientyst.com/to-json/ Last updated: 2022-08-28T09:13:12.000Z to\_json() ### How to Export DataFrame to JSON with Pandas URL: https://datascientyst.com/export-dataframe-to-json-pandas/ Last updated: 2023-03-18T09:50:39.000Z In this quick tutorial, we'll show how to **export DataFrame to JSON format in Pandas**. We will cover different export options. **(1) save DataFrame to a JSON file** ```python df.to_json('file.json') ``` **(2) change JSON format and data** ```python df.to_json('file.json', orient='split') ``` **Note:** Read also: [how to save Pandas DataFrame to JSON file without backslashes](https://datascientyst.com/convert-dataframe-to-json-without-backslash-in-pandas/) There are multiple options for parameter `orient` of method - [to\_json](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.to%5Fjson.html?ref=datascientyst.com): - `split` \- dict - `{‘index’ -> [index], ‘columns’ -> [columns], ‘data’ -> [values]}` - `records` \- list - `[{column -> value}, … , {column -> value}]` - `index` \- dict - `{column -> {index -> value}}` - `columns` \- dict (default for DataFrame) - `{column -> {index -> value}}` - `values` - just the values array - `table` \- dict - `{‘schema’: {schema}, ‘data’: {data}}` **Pro Tip 1** For Pandas Series several options are available for export to JSON {‘split’, ‘records’, ‘index’, ‘table’} ## Setup Let's create a sample DataFrame which will be exported to a JSON file: ```python import pandas as pd data = {'day': [1, 2, 3, 4, 5, 6, 7, 8], 'temp': [9, 8, 6, 13, 10, 15, 9, 10], 'humidity': [0.89, 0.86, 0.54, 0.73, 0.45, 0.63, 0.95, 0.67]} df = pd.DataFrame(data=data) ``` data looks like: | | day | temp | humidity | | - | --- | ---- | -------- | | 0 | 1 | 9 | 0.89 | | 1 | 2 | 8 | 0.86 | | 2 | 3 | 6 | 0.54 | | 3 | 4 | 13 | 0.73 | | 4 | 5 | 10 | 0.45 | ## 1: Export DataFrame as JSON file So let's start with the most simple example - exporting DataFrame as JSON string or a JSON file: ```python df.to_json('file.json') ``` If we provide the file name we will get a new JSON file created. The output file path can be provided as relative path: ```python df.to_json(r'/home/json/file.json') ``` **Pro Tip 1** NaN’s and None will be converted to null and datetime objects will be converted to UNIX timestamps. ## 2: Export DataFrame as JSON string Without file parameter - `path_or_buf` we will get JSON string: ```python df.to_json() ``` which will result into: ``` '{"day":{"0":1,"1":2,"2":3,"3":4,"4":5,"5":6,"6":7,"7":8},"temp":{"0":9,"1":8,"2":6,"3":13,"4":10,"5":15,"6":9,"7":10},"humidity":{"0":0.89,"1":0.86,"2":0.54,"3":0.73,"4":0.45,"5":0.63,"6":0.95,"7":0.67}}' ``` ## 3: Export DataFrame as pretty JSON Using parameter `indent=True` will **prettify the JSON output of Pandas** method `to_json()`: ```python print(df.to_json(indent=True)) ``` result: ``` { "day":{ "0":1, "1":2, "2":3, "3":4, "4":5, "5":6, "6":7, "7":8 }, "temp":{ "0":9, "1":8, "2":6, "3":13, "4":10, "5":15, "6":9, "7":10 }, "humidity":{ "0":0.89, "1":0.86, "2":0.54, "3":0.73, "4":0.45, "5":0.63, "6":0.95, "7":0.67 } } ``` ## 4: Export DataFrame JSON formats - `orient` ![](https://datascientyst.com/content/images/2022/08/export-dataframe-to-json-pandas.png) In this section we can see different JSON formats and data outputs. ### 4.1: columns - {column -> {index -> value}} Let's start by the default export option for DataFrame - `columns`: ```python df.to_json() ``` is equivalent to `df.to_json(orient='columns')` ``` {"day":{"0":1,"1":2,"2":3,"3":4,"4":5,"5":6,"6":7,"7":8}, "temp":{"0":9,"1":8,"2":6,"3":13,"4":10,"5":15,"6":9,"7":10}, "humidity":{"0":0.89,"1":0.86,"2":0.54,"3":0.73,"4":0.45,"5":0.63,"6":0.95,"7":0.67}} ``` ### 4.2: split - {‘index’ -> \[index\], ‘columns’ -> \[columns\], ‘data’ -> \[values\]} Using `orient` with option `split` will give us: ```python df.to_json(orient='split') ``` JSON with keys - `columns, index and data`: ``` {"columns":["day","temp","humidity"], "index":[0,1,2,3,4,5,6,7], "data":[[1,9,0.89],[2,8,0.86],[3,6,0.54],[4,13,0.73]..]} ``` ### 4.3: records - \[{column -> value}, … , {column -> value}\] For `records` we get list of dictionaries - **each row as a new entry in the output JSON**: ```python df.to_json(orient='records') ``` result: ``` [ {"day":1,"temp":9,"humidity":0.89}, {"day":2,"temp":8,"humidity":0.86}, {"day":3,"temp":6,"humidity":0.54}, {"day":4,"temp":13,"humidity":0.73}, {"day":5,"temp":10,"humidity":0.45}..] ``` ### 4.4: index - {index -> {column -> value}} With option `index` we got JSON file formatted as - index/data pairs: ```python df.to_json(orient='index') ``` result: ``` {"0":{"day":1,"temp":9,"humidity":0.89}, "1":{"day":2,"temp":8,"humidity":0.86}, "2":{"day":3,"temp":6,"humidity":0.54}, "3":{"day":4,"temp":13,"humidity":0.73}, "4":{"day":5,"temp":10,"humidity":0.45}..} ``` ### 4.5: orient='values' <-> df.values The option `orient='values'` is similar to `df.values`: ```python df.to_json(orient='values') ``` resulted JSON data is: ``` [[1,9,0.89],[2,8,0.86],[3,6,0.54],[4,13,0.73],[5,10,0.45],[6,15,0.63],[7,9,0.95],[8,10,0.67]] ``` ### 4.6: table - {‘schema’: {schema}, ‘data’: {data}} Finally we can use the option `table`. This give us **DB like and JSON schema like JSON export of a DataFrame**: ```python df.to_json(orient='table') ``` which results in next JSON schema: ``` { "schema": { "fields": [ { "name": "index", "type": "integer" }, { "name": "day", "type": "integer" }, { "name": "temp", "type": "integer" }, { "name": "humidity", "type": "number" } ], "primaryKey": [ "index" ], "pandas_version": "1.4.0" }, "data": [ { "index": 0, "day": 1, "temp": 9, "humidity": 0.89 }, { "index": 1, "day": 2, "temp": 8, "humidity": 0.86 }, { "index": 2, "day": 3, "temp": 6, "humidity": 0.54 }, { "index": 3, "day": 4, "temp": 13, "humidity": 0.73 } ] } ``` ## 5: Export DataFrame as JSON lines To save **Pandas DataFrame as JSON lines** we can use two parameters: - `orient='records'` - `lines=True` ```python df.to_json(orient='records', lines=True) ``` this gives JSON line output: ``` {"day":1,"temp":9,"humidity":0.89} {"day":2,"temp":8,"humidity":0.86} {"day":3,"temp":6,"humidity":0.54} {"day":4,"temp":13,"humidity":0.73} ``` ## Conclusion In this article we saw how to export DataFrame as a JSON string of files. We show how to pretty print the JSON output. We show different JSON formats and export to JSON lines. ### Create Evenly or Non-Evenly Spaced Series in Pandas and Python URL: https://datascientyst.com/create-evenly-non-evenly-spaced-series-pandas-python/ Last updated: 2022-08-26T20:47:51.000Z In this short tutorial, we'll see **how to create evenly or unevenly spaced arrays/series in Pandas and Python**. To create evenly spaced arrays we will use numpy methods: `np.linspace()` and `np.logspace()`. **(1) evenly spaced arrays/series** ```python np.linspace(1, 5) ``` **(2) non-evenly spaced arrays/series** ```python np.logspace(0, 1, 3) np.geomspace(1,3,5) ``` ## 1\. Create evenly spaced arrays in Python To create evenly spaced arrays in Python we can use the method: `np.linspace()`. It has 3 important parameters: - `start` \- The starting value of the sequence. - `stop` \- The end value of the sequence, unless endpoint is set to False. - `num` \- Number of samples to generate. Default is 50\. Must be non-negative. Some examples of evenly spaced arrays: ```python import numpy as np np.linspace(0, 5, 10) ``` result: ``` array([0. , 0.55555556, 1.11111111, 1.66666667, 2.22222222, 2.77777778, 3.33333333, 3.88888889, 4.44444444, 5. ]) ``` ```python np.linspace(1, 5) ``` result: ``` array([1. , 1.08163265, 1.16326531, 1.24489796, 1.32653061, 1.40816327, 1.48979592, 1.57142857, 1.65306122, 1.73469388, 1.81632653, 1.89795918, 1.97959184, 2.06122449, 2.14285714, 2.2244898 , 2.30612245, 2.3877551 , 2.46938776, 2.55102041, 2.63265306, 2.71428571, 2.79591837, 2.87755102, 2.95918367, 3.04081633, 3.12244898, 3.20408163, 3.28571429, 3.36734694, 3.44897959, 3.53061224, 3.6122449 , 3.69387755, 3.7755102 , 3.85714286, 3.93877551, 4.02040816, 4.10204082, 4.18367347, 4.26530612, 4.34693878, 4.42857143, 4.51020408, 4.59183673, 4.67346939, 4.75510204, 4.83673469, 4.91836735, 5. ]) ``` The official documentation is placed on: [numpy.linspace](https://numpy.org/doc/stable/reference/generated/numpy.linspace.html?ref=datascientyst.com) ## 2\. Create math functions with np.linspace() We can use `np.linspace()` to generate mathematical functions like: - sin - cos etc Below we can find visualization of math cos in Python and Matplotlib: ```python import matplotlib.pyplot as plt import numpy as np x = np.linspace( 0,10 ) for n in range(4): y = np.cos( x+n ) plt.figure(figsize=(12,5)) plt.title('chart:' + str(n)) plt.plot( x, y ) plt.show() ``` the result is plot of `cos` function: ![](https://datascientyst.com/content/images/2022/08/create-evenly-or-non-evenly-spaced-series-in-pandas.png) ## 3\. Create unevenly spaced arrays in Python **To make unevenly spaced arrays in Python and Pandas** we can use several numpy functions like: - `np.geomspace()` - `np.logspace()` Two examples of them: ```python np.geomspace(1,16,5) ``` will produce non evenly spaced array with 5 items from 1 to 16: ``` array([ 1., 2., 4., 8., 16.]) ``` While: ```python np.logspace(0, 2, 3) ``` will produce non evenly spaced array with 3 items from 1 to 100: ``` array([ 1., 10., 100.]) ``` In some cases the output is in scientific notation: ```python np.logspace(0, 4, 3) ``` result: ``` array([1.e+00, 1.e+02, 1.e+04]) ``` More info on: - [numpy.logspace](https://numpy.org/doc/stable/reference/generated/numpy.logspace.html?ref=datascientyst.com) - [numpy.geomspace](https://numpy.org/doc/stable/reference/generated/numpy.geomspace.html?ref=datascientyst.com) ## 4\. Make DataFrames with Uneven Array Lengths Finally let's cover a topic -\*\* how to create a DataFrame from non-even array length.\*\* If we try to create DataFrame from arrays with different length we will get an error: ```python import pandas as pd data = {"a": [1, 2, 3], "b": [1, 2, 3, 4]} pd.DataFrame(data) ``` the result is: ``` ValueError: All arrays must be of the same length ``` We can solve the error by method `from_dict` with parameter `orient='index'`: ```python pd.DataFrame.from_dict(data, orient='index').T ``` the result would be: | | a | b | | - | --- | --- | | 0 | 1.0 | 1.0 | | 1 | 2.0 | 2.0 | | 2 | 3.0 | 3.0 | | 3 | NaN | 4.0 | ## Conclusion We saw how to create evenly spaced arrays, simulate mathematical functions and visualize them. It was discussed how to generate non evenly spaced arrays by multiple functions. Finally we covered how to solve error - **ValueError: All arrays must be of the same length** \- and create DataFrame from non-evenly arrays lengths. ### How to Normalize Column or DataFrame in Pandas URL: https://datascientyst.com/normalize-column-pandas-dataframe/ Last updated: 2023-03-17T23:30:07.000Z In this tutorial, we'll learn **how to normalize columns or the whole DataFrame in Pandas**. We will show different ways like: **(1) Min Max normalization** for whole DataFrame ```python (df-df.min())/(df.max()-df.min()) ``` for column: ```python (df['col'] - df['col'].mean())/df['col'].std() ``` **(2) Mean normalization** ```python (df-df.mean())/df.std() ``` **(3) biased normalization** ```python scaler.fit_transform(df.iloc[:,:].to_numpy()) ``` Let's cover all examples in more detail. ## Setup For this post we are creating example DataFrame with 3 numeric columns: ```python import pandas as pd data = {'day': [1, 2, 3, 4, 5, 6, 7, 8], 'temp': [9, 8, 6, 13, 10, 15, 9, 10], 'humidity': [0.89, 0.86, 0.54, 0.73, 0.45, 0.63, 0.95, 0.67]} df = pd.DataFrame(data=data) ``` Data looks like: | | day | temp | humidity | | - | --- | ---- | -------- | | 0 | 1 | 9 | 0.89 | | 1 | 2 | 8 | 0.86 | | 2 | 3 | 6 | 0.54 | | 3 | 4 | 13 | 0.73 | | 4 | 5 | 10 | 0.45 | ## 1: Min Max normalization in Pandas So let's start by **min max normalization (called also min max scaling) in Pandas and Python**. ### Single column To do min max scaling for a single column we can do: ```python (df['humidity']-df['humidity'].min())/(df['humidity'].max()-df['humidity'].min()) ``` The result is normalized Series: ``` 0 0.88 1 0.82 2 0.18 3 0.56 4 0.00 5 0.36 6 1.00 7 0.44 Name: humidity, dtype: float64 ``` Checking data next to the original column: | | humidity\_norm | humidity | | - | -------------- | -------- | | 0 | 0.88 | 0.89 | | 1 | 0.82 | 0.86 | | 2 | 0.18 | 0.54 | | 3 | 0.56 | 0.73 | | 4 | 0.00 | 0.45 | ### All columns To **normalize all columns of a DataFrame** we can use: ```python (df-df.min())/(df.max()-df.min()) ``` Which will result into: | | day | temp | humidity | | - | -------- | -------- | -------- | | 0 | 0.000000 | 0.333333 | 0.88 | | 1 | 0.142857 | 0.222222 | 0.82 | | 2 | 0.285714 | 0.000000 | 0.18 | | 3 | 0.428571 | 0.777778 | 0.56 | | 4 | 0.571429 | 0.444444 | 0.00 | ## 2: Mean normalization in Pandas Next we can see how to do **mean normalization in Pandas and Python**. ### Single column For a single column we can apply mean normalization by: ```python (df['humidity'] - df['humidity'].mean())/df['humidity'].std() ``` The result and the original values: | | humidity\_norm | humidity | | - | -------------- | -------- | | 0 | 0.993475 | 0.89 | | 1 | 0.823165 | 0.86 | | 2 | \-0.993475 | 0.54 | | 3 | 0.085155 | 0.73 | | 4 | \-1.504406 | 0.45 | ### All columns To n**ormalize the whole DataFrame with mean normalization** we can do: ```python (df-df.mean())/df.std() ``` result: | | day | temp | humidity | | - | ---------- | ---------- | ---------- | | 0 | \-1.428869 | \-0.353553 | 0.993475 | | 1 | \-1.020621 | \-0.707107 | 0.823165 | | 2 | \-0.612372 | \-1.414214 | \-0.993475 | | 3 | \-0.204124 | 1.060660 | 0.085155 | | 4 | 0.204124 | 0.000000 | \-1.504406 | ![](https://datascientyst.com/content/images/2022/08/normalize-column-pandas-dataframe.png) ## 3: Biased normalization in Pandas To perform **biased normalization in Pandas** we can use the library `sklearn`. The results will differ from the Pandas normalization. ```python import pandas as pd from sklearn.preprocessing import StandardScaler scaler = StandardScaler() scaler.fit_transform(df.to_numpy()) ``` The results are: | | 0 | 1 | 2 | | - | ---------- | ---------- | ---------- | | 0 | \-1.527525 | \-0.377964 | 1.062070 | | 1 | \-1.091089 | \-0.755929 | 0.880001 | | 2 | \-0.654654 | \-1.511858 | \-1.062070 | | 3 | \-0.218218 | 1.133893 | 0.091035 | | 4 | 0.218218 | 0.000000 | \-1.608277 | ## 4: Normalize rows in Pandas There are multiple ways to **normalize rows**: - per sum - mean - min max ### Normalize rows by their sum To normalize row based on the sum of the row in Pandas we can do: ```python df.div(df.sum(axis=1), axis=0) ``` which will give use: | | day | temp | humidity | | - | -------- | -------- | -------- | | 0 | 0.091827 | 0.826446 | 0.081726 | | 1 | 0.184162 | 0.736648 | 0.079190 | | 2 | 0.314465 | 0.628931 | 0.056604 | | 3 | 0.225606 | 0.733221 | 0.041173 | | 4 | 0.323625 | 0.647249 | 0.029126 | ### Transpose To normalize row wise in Pandas we can combine: - `.T` to transpose rows to columns - `df.values` to get the values as numpy array Let's see an example: ```python import pandas as pd from sklearn import preprocessing data = df.T.values scaler = preprocessing.MinMaxScaler() pd.DataFrame(scaler.fit_transform(data)).T ``` So after using `df.values` we get: ``` array([[0.0135635 , 1. , 0. ], [0.15966387, 1. , 0. ], [0.45054945, 1. , 0. ], [0.26650367, 1. , 0. ], [0.47643979, 1. , 0. ], [0.3736952 , 1. , 0. ], [0.7515528 , 1. , 0. ], [0.78563773, 1. , 0. ]]) ``` which are transformed to: ``` array([[0. , 0.33333333, 0.88 ], [0.14285714, 0.22222222, 0.82 ], [0.28571429, 0. , 0.18 ], [0.42857143, 0.77777778, 0.56 ], [0.57142857, 0.44444444, 0. ], [0.71428571, 1. , 0.36 ], [0.85714286, 0.33333333, 1. ], [1. , 0.44444444, 0.44 ]]) ``` ## Conclusion In this article we learned how to normalize columns and DataFrame in Pandas. Different ways of normalization were covered like - biased, unbiased, normalization per sum. We also saw how to normalize rows of a DataFrame. Normalizing data is very useful in machine learning and visualizing data. ### How to Convert Decimal Comma to Decimal Point in Pandas DataFrame URL: https://datascientyst.com/convert-decimal-comma-to-decimal-point-pandas-dataframe/ Last updated: 2022-08-25T20:57:37.000Z In this quick tutorial, we're going to **convert decimal comma to decimal point in Pandas DataFrame** and vice versa. It will also show how to remove decimals from strings in Pandas columns. Different people in the world are using different decimal separator like: - decimal point - more often - decimal comma - in the Francophone area ## Setup Let's work with the following DataFrame: ```python import pandas as pd df = pd.DataFrame(data={'day': [1, 2, 3, 4, 5, 6, 7, 8], 'temp': [9, 8, 6, 13, 10, 15, 9, 10], 'humidity': [0.89, 0.86, 0.54, 0.73, 0.45, 0.63, 0.95, 0.67], 'humidity_eu': ['0,89', '0,86', '0,54', '0,73', '0,45', '0,63', '0,95', '0,67']}) ``` We have two columns with float data: - decimal comma - decimal point | | day | temp | humidity | humidity\_eu | | - | --- | ---- | -------- | ------------ | | 0 | 1 | 9 | 0.89 | 0,89 | | 1 | 2 | 8 | 0.86 | 0,86 | | 2 | 3 | 6 | 0.54 | 0,54 | | 3 | 4 | 13 | 0.73 | 0,73 | | 4 | 5 | 10 | 0.45 | 0,45 | ## 1: read\_csv - decimal point vs comma Let's start with the optimal solution - convert decimal comma to decimal point while reading CSV file in Pandas. Method `read_csv()` has parameter three parameters that can help: - `decimal` \- the decimal sign used in the CSV file - `delimiter` \- separator for the CSV file (tab, semi-colon etc) - `thousands` \- what is the symbol for thousands - if any To use them we can do: ```python df = pd.read_csv('file.csv', delimiter=";", decimal=",", thousands="`") ``` This will ensure that the correct decimal symbol is used for the DataFrame. ## 2: Convert comma to point If the DataFrame contains values with comma then we can convert them by `.str.replace()`: ```python df['humidity_eu'].str.replace(',', '.').astype(float) ``` result is: ``` 0 0.89 1 0.86 2 0.54 3 0.73 4 0.45 5 0.63 6 0.95 7 0.67 Name: humidity_eu, dtype: float64 ``` ## 3: Mixed decimal data - point and comma What can we do in case of mixed data in a given column? For this example we can use: list comprehensions and `pd.to_numeric()`. This can help us to identify the problematic values and keep the rest the same. For example we can do: ```python s = pd.Series(['0,89', '0,86', 0.54, 0.73, 0.45, '0,63', '0,95', '0,67']) mix = [x.replace(',', '.') if type(x) == str else x for x in s] ``` to replace the comma in all string records: ``` ['0.89', '0.86', 0.54, 0.73, 0.45, '0.63', '0.95', '0.67'] ``` Then we can convert the Series by: ```python pd.to_numeric(mix) ``` ``` array([0.89, 0.86, 0.54, 0.73, 0.45, 0.63, 0.95, 0.67]) ``` ## 4: Detect decimal comma in mixed column To detect which are the problematic values we can use: ```python s[pd.to_numeric(s, errors='coerce').isna() ] ``` the result of `to_numeric` is: ``` array([ nan, nan, 0.54, 0.73, 0.45, nan, nan, nan]) ``` while the final result is showing all values with decimal comma: ``` 0 0,89 1 0,86 5 0,63 6 0,95 7 0,67 dtype: object ``` ## 5: to\_csv - decimal point vs comma Finally if we like to write CSV file by method `to_csv` we can use parameters: - `decimal` - `sep` to control the decimal symbol. To convert CSV values from decimal comma to decimal point with Python and Pandas we can do : ```python df = pd.read_csv("file.csv", decimal=",") df.to_csv("test2.csv", sep=',', decimal='.') ``` ![](https://datascientyst.com/content/images/2022/08/convert-decimal-comma-to-decimal-point-pandas-dataframe.png) ## ValueError: could not convert string to float: '0,89' The error: "ValueError: could not convert string to float: '0,89'" is raised when we try to parse decimal comma to float. The error is given by method `s.astype(float)`: ```python s = pd.Series(['0,89', '0,86', 0.54, 0.73, 0.45, '0,63', '0,95', '0,67']) s.astype(float) ``` ## ValueError: Unable to parse string "0,89" at position 0 The error is the result of the `pd.to_numeric(s)` method - when a decimal comma is present in the input values. ```python s = pd.Series(['0,89', '0,86', 0.54, 0.73, 0.45, '0,63', '0,95', '0,67']) pd.to_numeric(s) ``` result: ``` ValueError: Unable to parse string "0,89" at position 0 ``` ## Conclusion In this article we saw how to replace, change and convert decimal symbols in Pandas. We saw how to detect problematic values in mixed columns - which have decimal commas and points simultaneously. Typical errors were explained. For further reference you can check also: - [How to Suppress and Format Scientific Notation in Pandas](https://datascientyst.com/suppress-format-scientific-notation-pandas/) - [How to Round Numbers in Pandas DataFrame](https://datascientyst.com/round-numbers-pandas-dataframe/) - [Solve - ValueError: could not convert string to float - Pandas](https://datascientyst.com/solve-valueerror-could-not-convert-string-to-float-pandas/) ### Timezone-aware and naive timestamp in Pandas URL: https://datascientyst.com/timezone-and-naive-timestamp-in-pandas/ Last updated: 2025-03-25T16:37:07.000Z In Pandas `datetime` can be two types: - `naive` \- no time zone information - `Timestamp('2022-08-24 16:53:23.186691')` - `aware` \- time zone information and local time - `Timestamp('2022-08-24 15:53:23.595457+0200', tz='Europe/Rome')` **Pro Tip 1** Always work with timezone information when possible. It's recommended to work with a single time zone - usually it's UTC. Below we can find several **Pandas examples of timestamps naive and aware**. ## Naive (local time) The first example shows how to **get the current date and time as a naive timestamp in Pandas.** ```python pd.Timestamp.now() ``` result: ``` Timestamp('2022-08-24 16:53:23.186691') ``` ## Timezone aware (UTC) Get **UTC time as timezone aware timestamp in Pandas:** ```python pd.Timestamp.utcnow() ``` result: ``` Timestamp('2022-08-24 13:53:23.387282+0000', tz='UTC') ``` ## Naive (local time) We can get the local time of any time zone as naive timestamp by: ```python pd.Timestamp.now(tz='Europe/Rome').tz_localize(None) ``` result: ``` Timestamp('2022-08-24 15:53:23.783935') ``` ## Timezone aware (local time) To use the time aware local time in Pandas: ```python pd.Timestamp.now(tz='Europe/Rome') ``` result: ``` Timestamp('2022-08-24 15:53:23.595457+0200', tz='Europe/Rome') ``` ## Naive (UTC) There are multiple ways to get naive UTC time. Below we will describe some of them. ### tz\_localize The method [pandas.DataFrame.tz\_localize](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.tz%5Flocalize.html?highlight=tz%5Flocalize&ref=datascientyst.com) will localize the tz-naive index of a Series or DataFrame to target the time zone. ```python pd.Timestamp.utcnow().tz_localize(None) ``` result: ``` Timestamp('2022-08-24 13:53:24.147123') ``` ### tz\_convert Method [pandas.DataFrame.tz\_convert](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.tz%5Fconvert.html?ref=datascientyst.com) convert the tz-aware axis to the target time zone. ```python pd.Timestamp.utcnow().tz_convert(None) ``` result: ``` Timestamp('2022-08-24 13:53:24.377683') ``` **Pro Tip 1** Difference between tz_convert and tz_localize is that first removes the timezone information resulting in naive local time while second remove the timezone information but converting to UTC, so giving naive UTC time ### tz\_convert and .now(tz='Europe/Rome') One more example for `tz_convert(None)`: ```python pd.Timestamp.now(tz='Europe/Rome').tz_convert(None) ``` result: ``` Timestamp('2022-08-24 13:53:23.983624') ``` ![](https://datascientyst.com/content/images/2022/08/timezone-and-naive-timestamp-in-pandas.png) ## Pandas to\_datetime - naive or aware Pandas `to_datetime` method has parameter: `utc`: - **True** \- the function always returns a timezone-aware UTC-localized Timestamp, Series or DatetimeIndex. - To do this, timezone-naive inputs are localized as UTC, while timezone-aware inputs are converted to UTC. - **False** \- inputs will not be coerced to UTC. - Timezone-naive inputs will remain naive, while timezone-aware ones will keep their time offsets. Limitations exist for mixed offsets (typically, daylight savings), see Examples section for details. More examples and explanation can be found here: [pandas.to\_datetime](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.to%5Fdatetime.html?ref=datascientyst.com) ### How to Convert Datetime to the Same Timezone in Pandas DataFrame URL: https://datascientyst.com/convert-datetime-the-same-timezone-pandas-dataframe/ Last updated: 2022-08-24T11:35:57.000Z In this post, we'll see how to convert `datetime` to the same timezone in Pandas DataFrame. You can also find how to solve errors like: - `ValueError: Tz-aware datetime.datetime cannot be converted to datetime64 unless utc=True` - `AttributeError: Can only use .dt accessor with datetimelike values` - `TypeError: Cannot localize tz-aware Timestamp, use tz_convert for conversions` So at the end you will get: 2015-05-11 03:00:00-04:00 -> 2016-10-07 08:30:19.428748+00:00 2016-10-07 08:30:19.428748+0000 -> 2015-05-11 07:00:00+00:00 or any other time zone. Below you can find the short answer: **(1) Convert the dates with utc=True** ```python df['date'] = pd.to_datetime(df['date'], utc=True) ``` **(2) Remove time zones** ```python df['date'].dt.tz_localize(None) ``` \*\*(3) apply + tz\_localize(None) \*\* ```python df['date'].apply(lambda x: pd.to_datetime(x).tz_localize(None)) ``` Let's cover all examples in the next section. ## Setup Suppose we have DataFrame like: ```python import pandas as pd dates = ['2016-10-07 08:30:19.428748+0000', '2016-10-07 08:30:19.428748+0000', '2015-05-11 03:00:00-04:00', '2015-05-10 22:00:00-05:00', '2015-05-10'] df = pd.DataFrame({'end_date': dates}) ``` data: | | end\_date | | - | ------------------------------- | | 0 | 2016-10-07 08:30:19.428748+0000 | | 1 | 2016-10-07 08:30:19.428748+0000 | | 2 | 2015-05-11 03:00:00-04:00 | | 3 | 2015-05-10 22:00:00-05:00 | | 4 | 2015-05-10 | ## 1: Convert datetime to the same time zone To convert datetime to the same time zone we can use method: `to_datetime()`: ```python df['date'] = pd.to_datetime(df['date'], utc=True) ``` After the conversion we will get: ``` 0 2016-10-07 08:30:19.428748+00:00 1 2016-10-07 08:30:19.428748+00:00 2 2015-05-11 07:00:00+00:00 3 2015-05-11 03:00:00+00:00 4 2015-05-10 00:00:00+00:00 Name: date, dtype: datetime64[ns, UTC] ``` As you can see the the time is changed for the records 2 and 3: `2015-05-10 22:00:00-05:00` \- > `2015-05-11 03:00:00+00:00` because of the change of the timezone. ## 2: Remove the time zone completely If you like to complete the timezone from your column or DataFrame we can do: ```python d = pd.Series(['2019-09-24 08:30:00-07:00', '2019-10-07 16:00:00-04:00', '2019-10-04 16:30:00+02:00']) d.apply(lambda x: pd.to_datetime(x).tz_localize(None)) ``` In the case above we need to use it because: - `pd.to_datetime(df['date'])` \- may raise error or give unexpected results - `df['time_tz'].dt.tz_localize(None)` \- the column is not datetime ## 3: Convert datetime with lambda + tz\_localize In some cases `pd.to_datetime()` will try to convert the dates but the column type will remain the object. For such cases we can use lambda for a custom conversion: ```python d = pd.Series(['2019-09-24 08:30:00-07:00', '2019-10-07 16:00:00-04:00', '2019-10-04 16:30:00+02:00']) d = pd.to_datetime(d) ``` The converted Series is from time object - which is not useful for us (check the errors described at the end): ``` 0 2019-09-24 08:30:00-07:00 1 2019-10-07 16:00:00-04:00 2 2019-10-04 16:30:00+02:00 dtype: object ``` To convert the column we can use combination of `.tz_localize(None)` and `apply`: ```python d.apply(lambda x: pd.to_datetime(x).tz_localize(None)) ``` The result is `datetime64`: ``` 0 2019-09-24 08:30:00 1 2019-10-07 16:00:00 2 2019-10-04 16:30:00 dtype: datetime64[ns] ``` ![](https://datascientyst.com/content/images/2022/08/convert-datetime-the-same-timezone-pandas-dataframe.png) ## ValueError: Tz-aware datetime.datetime cannot be converted to datetime64 unless utc=True Sometimes Pandas will not convert the column to datetime. In that case error: ``` ValueError: Tz-aware datetime.datetime cannot be converted to datetime64 unless utc=True ``` is raised. The code below will raise the error: ```python d = pd.Series(['2019-09-24 08:30:00-07:00', '2019-10-07 16:00:00-04:00', '2019-10-04 16:30:00+02:00']) d = pd.to_datetime(d) d = pd.to_datetime(d) ``` To solve the error and convert the column to date and not object we can use: ```python d.apply(lambda x: pd.to_datetime(x).tz_localize(None)) ``` ## AttributeError: Can only use .dt accessor with datetimelike values This error is shown when you try to use the '.dt' accessor on a non date column. In that case be sure that the column is from type date. We can simulate the error by: ```python pd.Series(['2000-10-01']).dt.date ``` and solve it with method: `pd.to_datetime()` ```python pd.to_datetime(pd.Series(['2000-10-01'])).dt.date ``` ## ValueError: Tz-aware datetime.datetime cannot be converted to datetime64 unless utc=True Finally let's cover the error: ``` ValueError: Tz-aware datetime.datetime cannot be converted to datetime64 unless utc=True ``` The error is returned from: ```python from datetime import datetime, timezone, timedelta d = datetime(2016, 5, 6, 12, tzinfo=timezone(-timedelta(hours=1))) pd.to_datetime(["2016-05-14 19:22 -0100", d]) ``` ## Conclusion We saw how to convert datetime and string to the same time zone in Pandas. Several errors typical for Pandas and date time were explained and solved. Finally we discussed alternative conversion when `pd.to_datetime()` is not working as expected. ### ValueError: Index contains duplicate entries, cannot reshape in Pandas URL: https://datascientyst.com/valueerror-index-contains-duplicate-entries-cannot-reshape-pandas/ Last updated: 2023-02-15T12:25:31.000Z In this tutorial, we'll see how to **solve a common Pandas error – "ValueError: Index contains duplicate entries, cannot reshape"**. We get this error from the Pandas when we try to reindex columns or index objects with duplicate entries. The error is very common for Pandas methods like: - `df.pivot()` - `df.unstack()` There are several quick solutions to it: **(1) Drop duplicates on the index or columns:** ```python df = df.drop_duplicates(['col_1','col_2']) ``` **(2) Use pivot\_table:** ```python pd.pivot_table(df, index='col_1', columns='col_2', values=['col_3'] , aggfunc='count') ``` **(3) Use reset the index** ```python .reset_index() ``` Let's cover several examples explaining the error in more context. ## Setup For this post we will work with the following DataFrame: ```python import pandas as pd data = { 'team': ['team_blue', 'team_red', 'team_blue', 'team_red', 'team_blue', 'team_red', 'team_blue', 'team_red', 'team_blue', 'team_red'], 'day': [7, 7, 8, 8, 9, 9, 10, 10, 10, 10], 'points': [10, 12, 9, 7, 14, 5, 9, 8, 10, 11], } df = pd.DataFrame(data) ``` data: | | team | day | points | | - | ---------- | --- | ------ | | 0 | team\_blue | 7 | 10 | | 1 | team\_red | 7 | 12 | | 2 | team\_blue | 8 | 9 | | 3 | team\_red | 8 | 7 | | 4 | team\_blue | 9 | 14 | ## Step 1: Detect duplicate records The error will be raised in case of duplicates on any of the axes. So in order to ensure that there's no duplicates on the selected columns or rows we can use: ```python df[df.duplicated(['team', 'day'], keep=False)] ``` output: | | team | day | points | | - | ---------- | --- | ------ | | 6 | team\_blue | 10 | 9 | | 7 | team\_red | 10 | 8 | | 8 | team\_blue | 10 | 10 | | 9 | team\_red | 10 | 11 | if we get non empty result it means that error will be occur if we try to use those columns with `pivot`: ```python pd.pivot(df,'team', 'day' )['points'] ``` The result is Pandas well known error: ``` ValueError: Index contains duplicate entries, cannot reshape ``` Note: We can use also: ```python df.duplicated(['row', 'col']).any() ``` The output is `True` or `False`. In case above it's `True` ## Step 2: Remove duplicate records To solve the error when we have duplicates we can remove duplicates. To do so we will use method: `df.drop_duplicates()`: ```python df = df.drop_duplicates(['team', 'day']) ``` This time running: ```python pd.pivot(df,'team', 'day' )['points'] ``` will not raise errors. The result is: | day | 7 | 8 | 9 | 10 | | ---------- | -- | - | -- | -- | | team | | | | | | team\_blue | 10 | 9 | 14 | 9 | | team\_red | 12 | 7 | 5 | 8 | ## Step 3: Use pivot\_table For duplicated records we can use the method `pivot_table` with aggregation functions. For example we can check number of records per each group by: ```python pd.pivot_table(df, index='day', columns='team', values=['points'] , aggfunc='count') ``` Instead of error we get the count for each column: | | points | | | ---- | ---------- | --------- | | team | team\_blue | team\_red | | day | | | | 7 | 1 | 1 | | 8 | 1 | 1 | | 9 | 1 | 1 | | 10 | 2 | 2 | **Pro Tip 1** Two main differences between: pivot and pivot_table (1) "pivot\_table" is a generalization of "pivot" that can handle duplicate values for index/column pair. (2) "pivot\_table" will only allow numeric types as "values=", whereas "pivot" will take string types. ## Step 4: Use reset\_index Depending on the data and the expected result we can use multiple transformations to solve the error. Let's demonstrate that with using method `reset_index()`: ```python df_agg = df.groupby(by=['team', 'day'])['points'].sum().reset_index() df_agg.pivot(index='team', columns='day', values=['points']) ``` The output is: | | points | | | | | ---------- | ------ | - | -- | -- | | day | 7 | 8 | 9 | 10 | | team | | | | | | team\_blue | 10 | 9 | 14 | 19 | | team\_red | 12 | 7 | 5 | 19 | How does it work? First we group by the two columns - `'team', 'day'` and then calculate the `sum()` of points for each group. ```python df.groupby(by=['team', 'day'])['points'].sum() ``` The result is a pandas Series: ``` team day team_blue 7 10 8 9 9 14 10 19 team_red 7 12 8 7 9 5 10 19 Name: points, dtype: int64 ``` **`reset_index()` convert it to a DataFrame without duplicates**: | | points | | | | | ---------- | ------ | - | -- | -- | | day | 7 | 8 | 9 | 10 | | team | | | | | | team\_blue | 10 | 9 | 14 | 19 | | team\_red | 12 | 7 | 5 | 19 | Now we can do `pivot()` without getting an error: "ValueError: Index contains duplicate entries, cannot reshape". **Pro Tip 2** The same result might be achieved in different ways. Depending on the context we can use different methods. ![](https://datascientyst.com/content/images/2022/08/valueerror-index-contains-duplicate-entries-cannot-reshape-pandas.png) ## Conclusion To sum up, this article shows how removing duplicates can solve the "ValueError: Index contains duplicate entries, cannot reshape" error. We also covered how to find the reasons for the error. Finally we show alternative ways to get the same result. ### Change Display Options of Pandas Styler by set_properties URL: https://datascientyst.com/change-display-options-pandas-styler-set_properties/ Last updated: 2022-08-23T09:00:43.000Z In this short post, we will see **how to use `set_properties` to change display options like**: - column width - color - column size - etc in Pandas DataFrame. We will show also methods like: - `set_table_attributes` - `set_table_styles` - `set_caption` ## Setup Suppose we have DataFrame like: ```python import pandas as pd data = { 'amount': [10.00, 20.5, 17.34], 'url': ['https://datascientyst.com/', 'https://datascientyst.com/pandas-most-typical-errors-and-solutions/', 'https://datascientyst.com/pandas-show-all-columns-rows/'], } df = pd.DataFrame(data) ``` with data(showing truncated data like Jupyter): | | amount | url | | - | ------ | ------------------------------------------------- | | 0 | 10.00 | https://datascientyst.com/ | | 1 | 20.50 | https://datascientyst.com/pandas-most-typical-... | | 2 | 17.34 | https://datascientyst.com/pandas-show-all-colu... | Displaying data in JupyterLab or Notebook will truncate the long strings by default. If we use styler for this DataFrame like: ```python df.style.background_gradient(cmap='Greens') ``` Then full width of the columns is displayed like: | | amount | url | | - | ------ | ------------------------------------------------------------------- | | 0 | 10.00 | https://datascientyst.com/ | | 1 | 20.50 | https://datascientyst.com/pandas-most-typical-errors-and-solutions/ | | 2 | 17.34 | https://datascientyst.com/pandas-show-all-columns-rows/ | ## 1\. Change display option of Styler Trying to change display options of Pandas like explained in: [How to show all columns and rows in Pandas](https://datascientyst.com/pandas-show-all-columns-rows/) \- by using `pd.set_option('display.max_rows', None)` will not work. In order to change **display options for Pandas styler** we can: - create new column with shorter width - use `.set_properties()` to change table properties like column width **Changing display option for a the whole DataFrame styler** ```python df.style.background_gradient(cmap='Greens').set_properties(**{'max-width': '200px'}) ``` or we can use **subset of column for which to change the table properties**: ```python df.style.background_gradient(cmap='Greens').set_properties(subset=['url'], **{'max-width': '200px'}) ``` ## 2\. Pandas set\_table\_attributes We can change more options by using method: `set_table_attributes`: ```python styler1.set_table_attributes('style="font-size: 25px"') ``` ## 3\. Pandas set\_table\_styles Another option is to use method: `set_table_styles` to change how Pandas styler is displayed in Jupyter: ```python styler1.set_table_styles([{'selector': '*', 'props': [('color', 'black'),('border-style','solid'),('border-width','1px')]}, {'selector': 'th', 'props': [('background-color', color1)]}]) ``` ## 4\. Combine multiple styles We can combine multiple methods at once to achieve nicer styling like: - Set DataFrame caption - change font size on DataFrame display ```python styler0.set_table_attributes("style='display:inline;font-size: 25px;'").set_caption('Before') ``` ## Conclusion In this article we saw how to change display options for Pandas Styler. We discussed multiple ways to change table properties for DataFrames. Finally we can achieve very beautiful DataFrame outlooks like shown on the image below: ![](https://datascientyst.com/content/images/2022/08/change-display-options-pandas-styler-set_properties.png) ### How to Iterate Through Multiple Rows at a Time in Pandas DataFrame URL: https://datascientyst.com/iterate-through-multiple-rows-time-pandas-dataframe/ Last updated: 2022-08-20T07:38:33.000Z In this short guide, I'll show you how to iterate simultaneously through 2 and more rows in Pandas DataFrame. So at the end you will get several rows into a single iteration of the Python loop. If you like to know more about more efficient way to iterate please check: [How to Iterate Over Rows in Pandas DataFrame](https://datascientyst.com/iterate-over-rows-pandas-dataframe/) ## Setup Let's create sample DataFrame to demonstrate iteration over multiple rows at once in Pandas: ```python import numpy as np import pandas as pd import string string.ascii_lowercase n = 5 m = 4 cols = string.ascii_lowercase[:m] df = pd.DataFrame(np.random.randint(0, n,size=(n , m)), columns=list(cols)) ``` Data will looks like: | | a | b | c | d | | - | - | - | - | - | | 0 | 1 | 1 | 2 | 2 | | 1 | 3 | 2 | 1 | 1 | | 2 | 2 | 2 | 3 | 4 | | 3 | 0 | 2 | 3 | 2 | | 4 | 0 | 4 | 3 | 4 | ## Step 1: Iterate over 2 rows - RangeIndex The most common example is to iterate over the default RangeIndex. To check if a DataFrame has `RangeIndex` or not we can use: ```python df.index ``` If the result is something like: ``` RangeIndex(start=0, stop=5, step=1) ``` Then we can use this method: ```python for i, g in df.groupby(df.index // 2): print(g) print('_' * 15) ``` This will give us: ``` a b c d 0 1 1 2 2 1 3 2 1 1 _______________ a b c d 2 2 2 3 4 3 0 2 3 2 _______________ a b c d 4 0 4 3 4 _______________ ``` To access the values inside the loop we can use: - row 1 - `g.values[0]` - row 2 - `g.values[1]` So we can see that the solution works fine even for an odd number of rows in DataFrame. How does it work? ```python df.index // 2 ``` will do mod on 2 resulting in: ``` Int64Index([0, 0, 1, 1, 2], dtype='int64') ``` Then we will group by the result `df.groupby(df.index // 2)` ![](https://datascientyst.com/content/images/2022/08/iterate-through-multiple-rows-time-pandas-dataframe.png) ## Step 2: Iterate over n rows at once So to iterate through n rows we need to change n in: `for i, g in df.groupby(df.index // n)`: ```python for i, g in df.groupby(df.index // 3): print(g) print('_' * 15) ``` result: ``` a b c d 0 1 1 2 2 1 3 2 1 1 2 2 2 3 4 _______________ a b c d 3 0 2 3 2 4 0 4 3 4 _______________ ``` ## Step 3: Iterate over Index A generic solution for DataFrame with non numeric index we can use numpy to split the index into groups like: ``` [0, 0, 1, 1, 2] ``` To do so we use method `np.arrange` providing the length of the DataFrame: ```python for i, g in df.groupby(np.arange(df.shape[0]) // 2): print (g) ``` result: ``` array([0, 0, 1, 1, 2]) ``` ## Step 4: Iterate with iterrows + zip Finally we can use `df.iterrows()` and `zip()` to iterate over multiple rows at once. The limitation of this method is that it doesn't work for an odd number of rows (last row is skipped). Also is not good for iterating over n rows. So combination of `df.iterrows()` and `zip()` to loop over 2 rows at the same time: ```python t = df.iterrows() for (i, r1), (j, r2) in zip(t, t): print(r1.values, r2.values) ``` result: ``` [1 1 2 2] [3 2 1 1] [2 2 3 4] [0 2 3 2] ``` ## Conclusion We saw how to loop over two and more rows at once in Pandas DataFrame. We covered the case of Index vs RangeIndex. Finally we saw an alternative way by combining `df.iterrows()` and `zip()` and the limitation of it. ### Convert string, K and M to number, Thousand and Million in Pandas/Python URL: https://datascientyst.com/convert-string-k-m-to-number-thousand-million-pandas-python/ Last updated: 2024-02-14T21:29:32.000Z In this short tutorial, we'll cover how to **convert natural language numerics like M and K into numbers with Pandas and Python**. We will show two different ways for **conversion of K and M to thousand and million**. We will also **cover the reverse case converting thousand and million to K and M**. So we will cover: - 0.1M to 100000 - 1 K - 1000 - 10000 to 10K - forty two - 42 Two Python libraries: - [humanize](https://pypi.org/project/humanize/?ref=datascientyst.com) \- turning a number into a fuzzy human-readable expression - [numerizer](https://pypi.org/project/numerizer/?ref=datascientyst.com) \- convert natural language numerics into ints and floats The image below show the examples: ![](https://datascientyst.com/content/images/2022/08/convert-k-m-to-thousand-million-pandas-python.png) ## Setup Let's use the following DataFrame the conversion from natural language numerals to numbers: ```python import pandas as pd import matplotlib.pyplot as plt data={'day': [1, 2, 3, 4, 5], 'numeric': [22, 222, '22K', '2M', '0.01 B'], 'numbers': [110, 11000, 1000000, 33300000, 456873], 'lang': ['one', 'five', 'twelve', 'forty two', 'one hundred and five']} df = pd.DataFrame(data, columns=['day', 'numeric', 'numbers', 'lang']) ``` Our data looks like: | | day | numeric | numbers | lang | | - | --- | ------- | -------- | -------------------- | | 0 | 1 | 22 | 110 | one | | 1 | 2 | 222 | 11000 | five | | 2 | 3 | 22K | 1000000 | twelve | | 3 | 4 | 2M | 33300000 | forty two | | 4 | 5 | 0.01 B | 456873 | one hundred and five | ## Step 1: Convert K/M to Thousand/Million First we will start the **conversion of large number abbreviations to numbers**. We will map the abbreviations to the math expression. So we will convert: ``` 22K -> 22 * 10**3 0.1 M -> 0.1 * 10**6 ``` ```python mp = {'K':' * 10**3', 'M':' * 10**6', 'B':' * 10**9', 't':' * 10**12', 'q':' * 10**15', 'Q':' * 10**15'} pd.eval(df['numeric'].replace(mp.keys(), mp.values(), regex=True)) ``` This will give us: ``` array([22.0, 222.0, 22000.0, 2000000.0, 10000000.0], dtype=object) ``` As we can see it works fine for columns with mixed data. It works well also for numeric values with spaces like `1 M` ### pd.eval limit 100 rows There seems to be limit of `pd.eval` and the returned results by `eval` are limited: ```python len(pd.eval([1 * 10**6] * 105)) ``` result: ``` 101 ``` Last five results are: ``` [1000000, 1000000, 1000000, 1000000, Ellipsis] ``` So the above code can be rewritten to: ```python df['cols_num'] = df['col'].replace(mp.keys(), mp.values(), regex=True) df['col_num'] = df.apply(lambda x: eval(x.col_num), axis=1) ``` to make it working for more than 100 rows. ## Step 2: Convert Thousand/Million to K/M For this step we will use the Python library - `humanize`. It will help us to\*\* convert easily and reliably large numbers to human readable abbreviations\*\*: ```python import humanize df['numbers'].apply(humanize.intword) ``` The result contains the converted values: ``` 0 110 1 11.0 thousand 2 1.0 million 3 33.3 million 4 456.9 thousand Name: numbers, dtype: object ``` The library supports multiple languages - about 25 like: - spanish - russian - french - portuguese Library can be installed by: ```bash pip install humanize ``` ## Step 3: Convert words to number - one to 1 Finally let's cover the case when we need to **convert language numerics into numbers**: - forty two -> 42 - twelve hundred -> 12000 - four hundred and sixty two -> 462 This time we will use Python library: numerize - which can be installed by: ```bash pip install numerize ``` So the Pandas code to convert numbers is: ```python from numerizer import numerize df['lang'].apply(numerize) ``` The output is: ``` 0 1 1 5 2 12 3 42 4 105 Name: lang, dtype: object ``` ## Large Number Abbreviations Below we can find a table of large number abbreviations which can be used to improve the mapping. | Abbreviation | Name | Value | Equivalent | | ------------ | --------------- | ------ | ---------- | | K | Thousand (Kilo) | 10^ 3 | 1000 | | M | Million | 10^ 6 | 1000K | | B | Billion | 10^ 9 | 1000M | | t | trillion | 10^ 12 | 1000B | | q | quadrillion | 10^ 15 | 1000t | | Q | Quintillion | 10^ 18 | 1000q | | s | sextillion | 10^ 21 | 1000Q | | S | Septillion | 10^ 24 | 1000s | | o | octillion | 10^ 27 | 1000S | | n | nonillion | 10^ 30 | 1000o | ## Conclusion In this post, we saw how to convert human readable and language expressions to numbers. We covered the reverse case of conversion of large numbers to abbreviations. It was shown how to map values to Python mathematical expressions and how to evaluate them. Two very useful Python libraries were used in Pandas for numeric conversion. ### NameError: name 'nan' is not defined in Pandas and Python URL: https://datascientyst.com/nan-in-mapper-name-nan-is-not-defined-pandas-python/ Last updated: 2022-08-17T10:21:47.000Z In this tutorial, we'll see how to solve a common Pandas error: > NameError: name 'nan' is not defined or > UndefinedVariableError: name 'nan' is not defined We get this error from the Pandas or Python when `numpy` is not imported as a library. ## 1\. import numpy nan By default numpy should be installed with Pandas. So to solve the error we need to import numpy `NaN` value we can use: ```python from numpy import nan ``` ×Note that NaN are aliases of nan . The error is common for map operation in Pandas. ```python mp = {'K':' * 10**3', 'M':' * 10**6', '< ': ''} df['amount'] = pd.eval(df['amount'].replace(mp.keys(), mp.values(), regex=True).str.replace(r'[^\d\.\*]+','')) ``` The example above will result into: > UndefinedVariableError: name 'nan' is not defined ## 2\. import numpy library As alternative we can import the whole library or other parts from `numpy`: ```python import numpy as np ``` This is described in: [10 minutes to pandas](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/10min.html?ref=datascientyst.com) ## 3\. install numpy Finally if you need to install numpy library you can do it by: ```python pip install numpy ``` ### Solve - ValueError: could not convert string to float - Pandas URL: https://datascientyst.com/solve-valueerror-could-not-convert-string-to-float-pandas/ Last updated: 2022-10-29T10:14:21.000Z In this short guide, I'll show you how to solve Pandas or Python errors: - `valueerror: could not convert string to float: '< "0.01'"` - `ValueError: could not convert string to float: '2,000'` - `ValueError: could not convert string to float: '$100.00'` - `ValueError: Unable to parse string "$10.00" at position 0` We will see how to solve the errors above and how to identify the problematic rows in Pandas. ## Setup Let's create an example DataFrame in order to reproduce the error: > ValueError: could not convert string to float: '$10.00' ```python import pandas as pd df = pd.DataFrame({'day': [1, 2, 3, 4, 5], 'amount': ['$10.00', '20.5', '17.34', '4,2', '111.00']}) ``` DataFrame looks like: | | day | amount | | - | --- | ------ | | 0 | 1 | $10.00 | | 1 | 2 | 20.5 | | 2 | 3 | 17.34 | | 3 | 4 | 4,2 | | 4 | 5 | 111.00 | ## Step 1: ValueError: could not convert string to float To convert string to float we can use the function: `.astype(float)`. If we try to do so for the column - amount: ```python df['amount'].astype(float) ``` we will face error: ``` ValueError: could not convert string to float: '$10.00' ``` ## Step 2: ValueError: Unable to parse string "$10.00" at position 0 We will see similar result if we try to convert the column to numerical by method `pd.to_numeric(`: ```python pd.to_numeric(df['amount']) ``` error is raised: ``` ValueError: Unable to parse string "$10.00" at position 0 ``` In both cases we can see the problematic value. But we are not sure if more different cases exist. ## Step 3: Identify problematic non numeric values To find which values are causing the problem we will use a simple trick. We will convert all values except the problematic ones by: ```python pd.to_numeric(df['amount'], errors='coerce') ``` This give us successfully converted float values: ``` 0 NaN 1 20.50 2 17.34 3 NaN 4 111.00 Name: amount, dtype: float64 ``` The ones in error will be replaced by `NaN`. No we can use mask to get only value which cause the error during the conversion to numerical values: ```python df[pd.to_numeric(df['amount'], errors='coerce').isna()]['amount'] ``` result is: ``` 0 $10.00 3 4,2 Name: amount, dtype: object ``` ![solve-valueerror-could-not-convert-string-to-float-pandas](https://datascientyst.com/content/images/2022/10/solve-valueerror-could-not-convert-string-to-float-pandas.webp) ## Step 4: Solve ValueError: could not convert string to float To solve the errors: > ValueError: could not convert string to float: '$10.00' > ValueError: Unable to parse string "$10.00" at position 0 We have several options: - ignore errors from invalid parsing and keep the output as it is: - `pd.to_numeric(df['amount'], errors='ignore')` - convert only numeric values and `NaN` for the rest in the column: - `pd.to_numeric(df['amount'], errors='coerce')` - fix problematic values In this step we are going to fix the problematic values. This can be done by replacing the non numeric symbols like: - `$` - `,` \- wrong decimal symbol etc So to replace the problematic characters we can use `str.replace`: ```python df['amount'].str.replace('$', '', regex=True) ``` Replacing multiple characters can be done by chaining multiple functions like.(for multiple replacing values): ```python df['amount'].str.replace('$', '', regex=True).replace('\,', '.', regex=True) ``` of by using regex (when all symbols are replaced by a single value): ```python df['amount'].str.replace('[$|,]', '', regex=True) ``` Finally we get only numeric values which can be converted to numeric column: 0 10.00 1 20.5 2 17.34 3 42 4 111.00 Name: amount, dtype: object ## Step 5: Convert numbers and keep the rest Finally if we like to convert only valid numbers we can use `errors='coerce'`. Then for all missing values we can populate them from the original column or Series: ```python data = pd.Series(['$10.00', '20.5', '17.34', '4,2', '111.00', np.NaN, '']) conv_data = pd.to_numeric(data, errors='coerce').fillna(data) conv_data ``` the result will be: ``` 0 $10.00 1 20.5 2 17.34 3 4,2 4 111.0 5 NaN 6 dtype: object ``` While using ```python pd.to_numeric(['$10.00', '20.5', '17.34', '4,2', '111.00', np.NaN, ''], errors='ignore') ``` will keep all the values the same if there is invalid parsing: ``` array(['$10.00', '20.5', '17.34', '4,2', '111.00', nan, ''], dtype=object) ``` ## Conclusion In this post, we saw how to properly convert strings to float columns in Pandas. We covered the most popular errors and how to solve them. Finally we discussed finding the problematic cases and fixing them. ### About DataScientYst URL: https://datascientyst.com/about-datascientyst/ Last updated: 2022-10-14T09:54:30.000Z ## Welcome to DataScientYst! ![welcome to datascientyst](https://datascientyst.com/content/images/2022/08/welcome_panda.png) Our mission is to simplify Data Science and make it accessible for everyone. We believe that complicated subjects don't have to be intimidating. Learning can be easy with the right information and maybe a bit more patience. You don't need formal education to learn and you can always learn something new. [![](https://datascientyst.com/content/images/2022/08/just_subscribe_to_datascientyst.png)](https://datascientyst.com/#/portal/signup) ## Is DataScientYst for you? ![](https://datascientyst.com/content/images/2022/08/pandas_circle.png)[DataScientYst.com](https://datascientyst.com/) was created to help people who would like to enter the vast world of Data Science.It could be perfect for you if you are: \- A student starting to learn Data Science or related subject \- A scientist who would like to use Data Science in his hobby/work \- Curious person who would like to understand our world better \- Data Scientist who would like to practice more and master the subject ## Why datascientYst? ![](https://datascientyst.com/content/images/2022/08/why_datascientyst.png)Exactly - because of the "Why?". It's our internal joke so let me explain - we changed the spelling of "scientist" to "scientYst" because "Y" represents the question "Why?". Data Science is about extracting knowledge and insights from data. Often knowledge is related to asking questions. Often Data Science projects starts by asking questions and answering them. So "Y" reminds us to stick to the origins. **Socrates** who is credited as the founder of Western philosophy created the Socratic method of questioning! **Did you know that** Socrates mentored Plato! Plato mentored Aristotle in philosophy! Aristotle mentored Alex the Great! Alex founded Alexandria....The Great Library of Alexandria The "Y" is also related to important skills for Data Scientists like: - curiositY - storY telling - strategY - TechnologY ## Our story ![datascientyst story](https://datascientyst.com/content/images/2022/08/datascientyst_creator.png) Hi, I'm Johny!I've started working with data back then in 2005 when I was a student in the university. I remember what I was thinking during a lecture for RDBMS - "Why should I care about this? I probably won't need it in the future... ". Nope - I was totally wrong about that. Several months later I've started job related to Oracle DB.My second jobs was also heavily related to DB, data validation, data migration.I've started blogging since 2017 on topics like Java, Python, SQL and Linux. I was writing articles and making videos. Initially it was started as a personal lab diary. Creating bread crumbs for my IT knowledge. But I've realized one thing – that I really like to help other people. Soon it become my passion to share knowledge and help people learn faster. My current job is also related to data - Data Quality - Collection - Analysis Yeah..I’ve made a lot of mistakes. I've learned a lot. So I would like to share that experience with other people. There is a lot to learn in Data Science, and **Data Science is definitely not easy. But it gets better with practice!** Subscribe to get future article, resource and updates! [![](https://datascientyst.com/content/images/2022/08/hi_subscribe_to_datascientyst.png)](https://datascientyst.com/#/portal/signup) To support us with small custom payments you can use: [Stripe - Support DataScientYst](https://buy.stripe.com/6oE15439hesJ848cMM?ref=datascientyst.com) ### How to Compare two Rows in Pandas URL: https://datascientyst.com/compare-two-rows-pandas/ Last updated: 2022-08-05T13:56:22.000Z In this short guide, we'll see how to **compare rows in Pandas DataFrame.** We will also cover how to find the difference between two rows in Pandas. To compare columns or DataFrames in Pandas please refer to: - [How to Compare Two Pandas DataFrames and Get Differences](https://datascientyst.com/compare-two-pandas-dataframes-get-differences/) - ## 2\. Setup In the post, we'll use the following DataFrame, which consists of several rows and columns: ```python import pandas as pd data = [('A',1, 0, 3, 1), ('A',1, 2, 5, 1), ('B',2, 1, 4, 3), ('B',3, 1, 0, 3), ('C',4, 3, 1, 2)] cols = ('col_1', 'col_2', 'col_3', 'col_4', 'col_5' ) df = pd.DataFrame(data, columns = cols) ``` Data is: | | col\_1 | col\_2 | col\_3 | col\_4 | col\_5 | | - | ------ | ------ | ------ | ------ | ------ | | 0 | A | 1 | 0 | 3 | 1 | | 1 | A | 1 | 2 | 5 | 1 | | 2 | B | 2 | 1 | 4 | 3 | | 3 | B | 3 | 1 | 0 | 3 | | 4 | C | 4 | 3 | 1 | 2 | ## Step 1: Compare two rows Pandas offers the method `compare()` which can be used in order of two rows in Pandas. Let's check how we can use it to compare specific rows in DataFrame. We are going to compare row with index - 0 to row - 2: ```python df.loc[0].compare(df.loc[2]) ``` The result is all values which has difference: | | self | other | | ------ | ---- | ----- | | col\_1 | A | B | | col\_2 | 1 | 2 | | col\_3 | 0 | 1 | | col\_4 | 3 | 4 | | col\_5 | 1 | 3 | Comparing the first two rows will give us only the differences: | | self | other | | ------ | ---- | ----- | | col\_3 | 0 | 2 | | col\_4 | 3 | 5 | ![](https://datascientyst.com/content/images/2022/08/compare-two-rows-pandas.png) ## Step 2: Compare two rows with highlighting To compare 2 rows while showing highlighted differences can be achieved with a custom function. First we will select the rows for comparison and transpose the resulting DataFrame by `df.loc[:1].T`. Then we will apply the function for comparison: ```python def highlight(d): df = pd.DataFrame(columns=d.columns, index=d.index) col1 = d.columns[0] col2 = d.columns[1] df[[col1, col2]] = 'background: None' df.loc[d[col1].ne(d[col2]), [col1, col2]] = 'background: yellow' return df df.loc[:1].T.style.apply(highlight, axis=None) ``` the result is: | | 0 | 1 | | ------ | - | - | | col\_1 | A | A | | col\_2 | 1 | 1 | | col\_3 | 0 | 2 | | col\_4 | 3 | 5 | | col\_5 | 1 | 1 | ## Step 3: Calculate the difference between rows To calculate the difference between rows - row by row we can use the function - `diff()`. To calculated difference for `col_4`: ```python df['col_4'].diff() ``` the result would be difference calculated for each two rows: ``` 0 NaN 1 2.0 2 -1.0 3 -4.0 4 1.0 Name: col_4, dtype: float64 ``` Method `diff()` has two useful parameters: - `periods` \- shift of periods for calculating difference - `axis` \- rows or columns If there are string values in some rows you may face the following error: > TypeError: unsupported operand type(s) for -: 'str' and 'str' ## Step 4: Compare rows from different DataFrames In order to compare rows from different DataFrames we can select them by index: ```python df2 = df.copy() df.loc[0].compare(df2.loc[1]) ``` Another option is to match rows on another column - let say unique column - `url`. We can match the records by: ```python url_same = df1[df1.url.isin(df2.url)].url ``` Then we can select one by: ```python url = url_same[2] ``` Finally we can concatenated into single DataFrame as: ```python pd.concat([df1[df1.url == url], df2[df2.url == url]]).T ``` Finally we can transpose the result for readability. The full code: ```python url_same = df1[df1.url.isin(df2.url)].url url = url_same[2] df_c = pd.concat([df1[df1.url == url], df2[df2.url == url]]).T ``` ## Step 5: Finding the duplicate rows We can find the duplicated rows in the whole DataFrame by command like: ```python df[['col_1', 'col_2']].duplicated() ``` Method `duplicated()` returns boolean array for duplicate rows: ``` 0 False 1 True 2 False 3 False 4 False dtype: bool ``` or for the whole DataFrame. In this case we will get also the data: ```python df[df.duplicated()] ``` Empty DataFrame since there no duplicated rows on every column ## Step 6: Finding the unique rows To find all the unique rows we can use: ```python df.drop_duplicates() ``` or for subset of columns: ```python df[['col_1', 'col_2']].drop_duplicates() ``` result: | | col\_1 | col\_2 | | - | ------ | ------ | | 0 | A | 1 | | 2 | B | 2 | | 3 | B | 3 | | 4 | C | 4 | ## Conclusion In this article we show how to compare rows, compare rows with highlighting and extract the difference of two rows. We covered comparison of rows from different DataFrames and calculating the difference between rows for the whole DataFrame. ### Get value_counts for Multiple Columns in Pandas URL: https://datascientyst.com/get-value_counts-for-multiple-columns-in-pandas/ Last updated: 2022-08-05T08:39:01.000Z Need to **get value\_counts for multiple columns in Pandas DataFrame**? In this article we will cover several options to get value counts for multiple columns or the whole DatFrame. ## Setup Let's create a sample DataFrame which will be used in the next examples: ```python import numpy as np import pandas as pd df = pd.DataFrame(np.random.randint(0, 2, (5, 3)), columns=list('ABC')) ``` The tabular representation of the DataFrame is: | | A | B | C | | - | - | - | - | | 0 | 1 | 0 | 1 | | 1 | 0 | 1 | 1 | | 2 | 0 | 1 | 1 | | 3 | 0 | 0 | 1 | | 4 | 0 | 1 | 1 | ## Step 1: Apply value\_counts on several columns Let's start with applying the function `value_counts()` on several columns. This can be done by using the function `apply()`. We can list the columns of our interest: ```python df[['A', 'B']].apply(pd.value_counts) ``` The results is a DataFrame with the count for columns: | | A | B | | - | - | - | | 0 | 4 | 2 | | 1 | 1 | 3 | ## Step 2: Apply value\_counts with parameters What if we like to use `value_counts()` on multiple columns with parameters? Then we can pass them the the `apply()` function as: ```python df[['A', 'B']].apply(pd.value_counts, normalize=True) ``` The result is the normalized count of columns A and B: | | A | B | | - | --- | --- | | 0 | 0.8 | 0.4 | | 1 | 0.2 | 0.6 | ![](https://datascientyst.com/content/images/2022/08/get-value_counts-for-multiple-columns-in-pandas.png) ## Step 3: Apply value\_counts on all columns To apply `value_counts()` on every column in a DataFrame we can use the same syntax as before: ```python df.apply(pd.value_counts) ``` The result is the count of each column: | | A | B | C | | - | - | - | --- | | 0 | 4 | 2 | NaN | | 1 | 1 | 3 | 5.0 | ## Step 4: Simulate value\_counts with melt Finally let's check how we can use advanced analytics in order to manipulate data. We will use the `melt()` function in order to reshape the original DataFrame and get count for columns. The first step is to use `melt()`: ```python df.melt(var_name='column', value_name='value') ``` This will change data into two columns - in other words - DataFrame will be represented row wise in columns: - column - the source column - value - the value of the column | | column | value | | - | ------ | ----- | | 0 | A | 1 | | 1 | A | 0 | | 2 | A | 0 | | 3 | A | 0 | | 4 | A | 0 | Now we can apply `value_counts()`: ```python df.melt(var_name='column', value_name='value').value_counts() ``` To get result as: ``` column value C 1 5 A 0 4 B 1 3 0 2 A 1 1 dtype: int64 ``` or we can display the results as a DataFrame with sorted counts: ```python (pd.DataFrame( df.melt(var_name='column', value_name='value').value_counts()) .sort_values(by=['column']).rename(columns={0: 'count'})) ``` | | | count | | ------ | ----- | ----- | | column | value | | | A | 0 | 4 | | 1 | 1 | | | B | 1 | 3 | | 0 | 2 | | | C | 1 | 5 | × **Pro Tip 1** **Advanced users** can go further and combine: pd.crosstab() and df.melt ```python pd.crosstab(**df.melt(var_name='columns', value_name='index')) ``` This will result into: | columns | A | B | C | | ------- | - | - | - | | index | | | | | 0 | 4 | 2 | 0 | | 1 | 1 | 3 | 5 | ## Conclusion To summarize we saw how to apply `value_counts()` on multiple columns. We covered how to use `value_counts()` with parameters and for every column in DataFrame. Finally we discussed advanced data analytics techniques to get count of value for multiple columns. ### How to Compare Two Columns in Pandas (Highlight) URL: https://datascientyst.com/compare-two-columns-highlight-pandas/ Last updated: 2023-03-17T23:06:44.000Z In this article, we will see how to **compare (with highlight of differences) two columns in Pandas DataFrame**. If you like to see how to compare two DataFrames in Pandas please check: [How to Compare Two Pandas DataFrames and Get Differences](https://datascientyst.com/compare-two-pandas-dataframes-get-differences/) ## Setup In the post, we'll use the following DataFrame, which consists of several rows and columns: ```python import pandas as pd data = {'col1': ['apple', 'pear', 'apple', 'orange', 'lemon'], 'col2': ['apple', 'pear', 'lemon', 'orange', 'apple']} df = pd.DataFrame(data, columns = ['col1', 'col2']) ``` DataFrame looks like: | | col1 | col2 | | - | ------ | ------ | | 0 | apple | apple | | 1 | pear | pear | | 2 | apple | lemon | | 3 | orange | orange | | 4 | lemon | apple | ## Step 1: Compare with map and lambda The first way to **compare the two columns of a DataFrame** is by chaining: - `map` - `lambda` ```python df.style.apply(lambda x: (x != df['col1']).map({True: 'background-color: yellow', False: ''}), subset=['col2']) ``` We need to enter the column names. After the execution the differences between the two columns will be highlighted in yellow: | | col1 | col2 | | - | ------ | ------ | | 0 | apple | apple | | 1 | pear | pear | | 2 | apple | lemon | | 3 | orange | orange | | 4 | lemon | apple | ## Step 2: Compare and highlight by custom function In this step we are going to **define a function to compare the first two columns of a DataFrame**. After that we are going to highlight the row which has a difference in the values. The columns will be chosen by index: `d.columns[0]` ```python def highlight_diff(d): df = pd.DataFrame(columns=d.columns, index=d.index) col1 = d.columns[0] col2 = d.columns[1] df[[col1, col2]] = 'background: None' df.loc[d[col1].ne(d[col2]), [col1, col2]] = 'background: yellow' return df df.style.apply(highlight_diff, axis=None) ``` The resulted DataFrame contains highlighted rows where information is different for the compared columns: | | col1 | col2 | | - | ------ | ------ | | 0 | apple | apple | | 1 | pear | pear | | 2 | apple | lemon | | 3 | orange | orange | | 4 | lemon | apple | ## Step 3\. Compare and get difference If we want to **compare two columns and get only the different rows** we can use method: `compare()`: ```python df['col1'].compare(df['col2']) ``` In that case we will get only the rows which has some difference: | | self | other | | - | ----- | ----- | | 2 | apple | lemon | | 4 | lemon | apple | ## Step 4\. Compare columns of different DataFrames If we need to **compare columns from different DataFrames** like the example below: ```python import pandas as pd data = {'col1': ['apple', 'pear', 'apple', 'orange', 'lemon'], 'col2': ['apple', 'pear', 'lemon', 'orange', 'apple']} df1 = pd.DataFrame(data, columns = ['col1']) df2 = pd.DataFrame(data, columns = ['col2']) ``` we can do direct comparison by: ```python df1['col1'] != df2['col2'] ``` which will give us: ``` 0 False 1 False 2 True 3 False 4 True dtype: bool ``` or we can concatenate them into a new DataFrame: ```python pd.concat([df1['col1'], df2['col2']], axis=1) ``` and then apply the previous steps. ## Step 5\. Compare multiple columns Let's see how we can **compare multiple columns** at the same time from one DataFrame. Let's work with the following DataFrame: ```python import pandas as pd data = {'col1': ['apple', 'pear', 'apple', 'orange', 'lemon'], 'col2': ['apple', 'pear', 'lemon', 'orange', 'apple'], 'col3': [1,1,0,1,0], 'col4': [1,1,0,1,0]} df = pd.DataFrame(data, columns = ['col1', 'col3', 'col2', 'col4']) ``` data looks like: | | col1 | col3 | col2 | col4 | | - | ------ | ---- | ------ | ---- | | 0 | apple | 1 | apple | 1 | | 1 | pear | 1 | pear | 1 | | 2 | apple | 0 | lemon | 0 | | 3 | orange | 1 | orange | 1 | | 4 | lemon | 0 | apple | 0 | To compare the first two columns with the last two columns we can use broadcasting with method `all()`: ```python (df.iloc[:, 0:2].values == (df.iloc[:, 2:4].values)).all(axis=1) ``` This will give us: ``` array([ True, True, False, True, False]) ``` Showing that the first two rows have identical values for that comparison. Depending on the context we can use this method to compare multiple columns with different operators like: - `>` \- greater - `-` \- subtraction - different column numbers ## Conclusion In this article we saw how to compare two columns in Pandas. We covered highlighting the differences or just extracting them. We also covered how to compare multiple columns or columns from different DataFrames. It was also a way for conditional comparison of columns. ### How to Round Numbers in Pandas DataFrame URL: https://datascientyst.com/round-numbers-pandas-dataframe/ Last updated: 2022-08-01T07:06:39.000Z In this tutorial, we'll see how to **round float values in Pandas.** We'll illustrate several examples related to rounding numbers in DataFrame like: - round number to 2 decimal places - round up - round down - round float to int - round number to nearest - round single column - round whole DataFrame If you need to format or suppress scientific notation in Pandas please check: [How to Suppress and Format Scientific Notation in Pandas](https://datascientyst.com/suppress-format-scientific-notation-pandas/) ## Setup For this article we are going to create Pandas DataFrame with random float numbers. To generate random float numbers in Python and Pandas we can use the following code: Generating N random float numbers in Python: ```python import random nums = [] for i in range(10): x = random.uniform(0.1, 10.0) nums.append(round(x, 5)) ``` Next we will create DataFrame from the above data: ```python import pandas as pd data = {'val1': nums[:5], 'val2': nums[5:]} df = pd.DataFrame(data, columns = ['val1', 'val2']) ``` DataFrame looks like: | | val1 | val2 | | - | ------- | ------- | | 0 | 3.73156 | 3.62060 | | 1 | 0.34162 | 0.93644 | | 2 | 1.25441 | 9.39390 | | 3 | 3.65596 | 2.66972 | | 4 | 3.38249 | 6.26096 | ## Step 1: Round up values to 2 decimal places Let's start by **rounding specific column up to 2 decimal places**. ```python df['val1'].round(decimals = 2) ``` or simply: ```python df['val1'].round(2) ``` The result is a Series from the rounded data: ``` 0 3.73 1 0.34 2 1.25 3 3.66 4 3.38 Name: val1, dtype: float64 ``` So we can see that data was rounded up. For example: ``` 3.65596 ``` become ``` 3.66 ``` More information about method round is available here: [DataFrame.round](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.round.html?ref=datascientyst.com) ## Step 2: Round down numbers in specific column Next let's cover how to **round down a column in Python and Pandas**. We will use `np.floor` for rounding down: ```python import numpy as np df['val1'].apply(np.floor) ``` result is: ``` 0 3.0 1 0.0 2 1.0 3 3.0 4 3.0 Name: val1, dtype: float64 ``` If you need decimal precision or custom rounding down please refer to step 4. ## Step 3: Round up values We can use `np.ceil` to round up numbers in Pandas: ```python import numpy as np df['val1'].apply(np.ceil) ``` result is: ``` 0 4.0 1 1.0 2 2.0 3 4.0 4 4.0 Name: val1, dtype: float64 ``` If you need decimal precision or custom rounding up please refer to the next step. ![](https://images.unsplash.com/photo-1529078155058-5d716f45d604?crop=entropy&cs=tinysrgb&fit=max&fm=jpg&ixid=MnwxMTc3M3wwfDF8c2VhcmNofDR8fG51bWJlcnxlbnwwfHx8fDE2NTkzMzc1NTI&ixlib=rb-1.2.1&q=80&w=2000) ## Step 4: Round down values with decimal precision We can define **custom function in order to round down values in Pandas with decimal precision**. This allow us to use any custom logic to round numbers in Pandas: ```python def my_floor(a, precision=0): return np.true_divide(np.floor(a * 10**precision), 10**precision) df['val1'].apply(my_floor, precision=3) ``` The result is round down float to the 3rd decimal place: ``` 0 3.731 1 0.341 2 1.254 3 3.655 4 3.382 Name: val1, dtype: float64 ``` ## Step 5: Round number to nearest integer We can **round floats to nearest integer in Pandas** by combining: - `round()` - `astype()` The code is: ```python df['val1'].round(0).astype(int) ``` and the result is: ``` 0 4 1 0 2 1 3 4 4 3 Name: val1, dtype: int64 ``` ## Step 6: Round values in the whole DataFrame So far we've worked with single columns. If you need to apply **rounding operations over the whole DataFrame** we can use method `round()`: ```python df.round(3) ``` The result is rounding up to 3 decimal places: | | val1 | val2 | | - | ----- | ----- | | 0 | 3.732 | 3.621 | | 1 | 0.342 | 0.936 | | 2 | 1.254 | 9.394 | | 3 | 3.656 | 2.670 | | 4 | 3.382 | 6.261 | ## Conclusion To summarize, in this article, we've seen examples of rounding numbers and values in Pandas. We briefly described rounding up and down. And finally, we've seen how to apply rounding on specific columns or the whole DataFrame. ### First steps in Data Science (discussion) URL: https://datascientyst.com/first-steps-in-data-science-discussion/ Last updated: 2022-08-01T05:55:13.000Z This post is taken from a discussion in a Slack group related to **first steps in Data Science**: [#data\_science in Python Developers](https://pythondev.slack.com/archives/C0JB9ATQV?ref=datascientyst.com). Only a few small edits were done from my side. There is valuable information which might help beginners in Data Science. There is also advice on learning in general. Hope that will help you when you start your journey in Data Science! ## Question 1 - First steps in data science? *Hello, I am a 2nd year IT student. Сan you help me with my first steps in data science? Please tell me what resources to use* ### Advice 1 - In the beginning If you're at the very start I would probably start by getting a book and following it, and maybe Google blog posts as well. Example Search: [https://www.google.com/search?channel=fs&client=ubuntu&q=best+datascience+books](https://www.google.com/search?channel=fs&client=ubuntu&q=best+datascience+books&ref=datascientyst.com) Example Data Science Book(from me):[Python Data Science Handbook](https://jakevdp.github.io/PythonDataScienceHandbook/?ref=datascientyst.com) × **Pro Tip 1** If you're at the very start I would probably start by getting a book ## Question 2 - Where to get the practice in Data Science *Where to get the practice. If the theory is clear, can you recommend some project?* ### Advice 2 - Choose area You can try out some kaggle competitions or something to start. Are you interested in learning ML, or more analytics focus? If you are not sure about the above read the intro chapter of **Book:** [Data Science in Context](https://datascienceincontext.com/?ref=datascientyst.com) × **Pro Tip 2** If you are not sure how to start read the intro chapter of **Data Science in Context** ### Advice 3 - The process from end to end To start in data science ... you'd need to **study linear algebra & statistics for the basic mathematical foundation.** Then, for the technical skills, you have to learn the **basic tools - Python, R, and various toolkits that these include (such as Spark, etc)**. Then you need to acquire the ability to put these items into good use. **Thinking is more important than copypasting what someone else did on their Medium blog.** It's a long programme of study, if you are just starting × **Pro Tip 3** Thinking is more important than copypasting what someone else did on their Medium blog. There is a vague line that divides data science from data analysis, and I am pretty sure that the line is defined by your mathematical abilities. Now, having said that, most places will not have a business need that goes beyond an activity that is more properly defined as data analysis, so it is instructive to realize that data science is a poorly defined thing. In a nutshell - **get good at Python & R**, and start applying to jobs with proof that you know both. Just let the cookie crumble at that point. oh, I nearly forgot - **SQL is a must.** ### Advice 4 - Interest as motivation If I am interested in a subject, going for the core materials is insightful. If not, **I learn better by treating a subject as a black box that solves the problem I actually care about, and then familiarizing myself gradually and organically with the topic as I use a tool more and more, without ever having studied the "basics"**. Due to that I studied a lot of Linear Algebra / Combinatorics as I find them very interesting topics, but I have used most software ( Keycloak, Airflow, RabbitMQ,... ) as tools that get the job I need done, without ever looking into the core CS topics. My point being that if Linear Algebra doesn't appeal to you directly... I might start with tutorials that show directly use cases: - Word Vec with NLP - Image Labeling with Computer Vision to keep the motivation up, and build back from it. It comes down to what you find more interesting, the problem being solved, or the technologies used to solve it. ![](https://images.unsplash.com/photo-1532619675605-1ede6c2ed2b0?crop=entropy&cs=tinysrgb&fit=max&fm=jpg&ixid=MnwxMTc3M3wwfDF8c2VhcmNofDQzfHxkYXRhJTIwc2NpZW5jZXxlbnwwfHx8fDE2NTkzMzI3MTI&ixlib=rb-1.2.1&q=80&w=2000) ### Advice 5 - top-down vs bottom-up learning There are 2 ways to learn: - top-down - bottom-up. **Top-down is preferable if you're a professional that lacks the time to spend internalizing basic concepts in a typical academic environment.** My advice to people in this situation is to monkey-see : monkey-do until you magically "get it". In other words, top-down. In the military they would call this "results oriented training." So I suggest **learning by example, and taking the time to understand the fundamentals as you go through each case.** Still takes a long time, but much more effective IMO × **Pro Tip 4** A single example, if well chosen, can produce an incredibly productive lesson You could, for instance, go through a **well-chosen deep learning example**, of which there is a myriad, and use it to understand weights, optimization algorithms (stochastic gradient descent, etc), the meaning of under / overtraining, training set selection strategies, neural net architectures, and so on. **A single example, if well chosen, can produce an incredibly productive lesson**. You can consider yourself graduated if you can cogently interpret a publication in the field, and implement it in practice. ### Advice 6 - Understand data. Master 1 thing Usually when I am learning I do both ways for me it is easier to get and know fundamentals. (it is plenty of time if you manage it well - drop tv watching, youtube, disable all notifications on the phone, delete games if that distracts you). Try to do what others have done and figure out each step what they have done. **Get some books, articles, videos about it and learn from it**. Check code libraries documentation (at the end of the day you need to use some code it is very rare that you will write everything yourself), knowing what each library can give you is powerful. Looking into others codes helps to figure out structure and see if it fits what you want or you can put it more clearly with better design patterns ( this is form CS degree hitting). × **Pro Tip 5** understand what actual data means and how it was captured **Concentrate on 1 thing and learn it very well if you are really willing to get into that field.** Knowing a single thing but very well helps to figure out other concepts, usually specific areas always have something in common. I think that in data science is important to really **understand what actual data means and how it was captured**. ## Question 3 - Where to get the practice in Data Science *I am pursuing for Data Science in my college degree* *Any tip on what all should I do from the start of my college to be good at DS* *also, i don't have any previous knowledge of programming* ### Advice 7 Learn basics of Python and Python notebook + matplotlib library ([https://matplotlib.org/](https://matplotlib.org/?ref=datascientyst.com)) Learn about different data storages and databases ( SQL, no-SQL queries), merging data, general statistics. Take a public dataset and try to understand what data is showing. ### How to Normalize JSON or Dict to New Columns in Pandas URL: https://datascientyst.com/normalize-json-dict-new-columns-pandas/ Last updated: 2024-04-09T10:14:37.000Z In this article, we will see how to **convert JSON or string representation of dictionaries in Pandas.** JSON(JavaScript Object Notation) data and dictionaries can be stored and imported in different ways. This might result in unexpected results or need to convert them to new columns. In order to convert **JSON, dicts and lists to tabular form** we can use several different options. Let's cover the most popular of them in next steps. ## Setup In the post, we'll use the following DataFrame, which has columns: - `col_json` \- JSON column stored as JSON - `col_str_dict` \- JSON column stored as a string - `col_str_dict_list` \- JSON column with nested list ```python import pandas as pd data = {'col_json': {0: {'x': 1, 'y': 0, 'xy':1}, 1: {'x': 0, 'y': 1, 'xy':1}, 2: {'x': 1, 'y': 1, 'xy':1}}, 'col_str_dict': {0: "{'x': 1, 'y': 0, 'xy':1}", 1: "{'x': 0, 'y': 1, 'xy':1}", 2: "{'x': 1, 'y': 1, 'xy':1}"} , 'col_str_dict_list': {0: '{"x": [1,1]}', 1: '{"x": [0,1]}', 2: '{"x": [1,0]}'} } df = pd.DataFrame(data) df ``` It's not possible to distinguish how JSON is stored from the dtypes: ```python df.dtypes ``` results into: ``` col_json object col_str_dict object col_str_dict_list object dtype: object ``` DataFrame data looks like: | | col\_json | col\_str\_dict | col\_str\_dict\_list | | - | ------------------------- | ------------------------ | -------------------- | | 0 | {'x': 1, 'y': 0, 'xy': 1} | {'x': 1, 'y': 0, 'xy':1} | {"x": \[1,1\]} | | 1 | {'x': 0, 'y': 1, 'xy': 1} | {'x': 0, 'y': 1, 'xy':1} | {"x": \[0,1\]} | | 2 | {'x': 1, 'y': 1, 'xy': 1} | {'x': 1, 'y': 1, 'xy':1} | {"x": \[1,0\]} | ## 1: Normalize JSON - json\_normalize Since Pandas version 1.2.4 there is new method to normalize JSON data: `pd.json_normalize()` It can be used to **convert a JSON column to multiple columns**: ```python pd.json_normalize(df['col_json']) ``` this will result into new DataFrame with values stored in the JSON: | | x | y | xy | | - | - | - | -- | | 0 | 1 | 0 | 1 | | 1 | 0 | 1 | 1 | | 2 | 1 | 1 | 1 | The method [pd.json\_normalize](https://pandas.pydata.org/docs/reference/api/pandas.json%5Fnormalize.html?ref=datascientyst.com) has several parameters like: - `record_path` \- Path in each object to list of records. If not passed, data will be assumed to be an array of records. - `meta` \- Fields to use as metadata for each record in the resulting table. - `errors` \- Configures error handling - `ignore` or `raise` - `max_level` \- Max number of levels(depth of dict) to normalize. if None, normalizes all levels. How to use them in more complex examples is covered in Step 5. ## 2: Flattening JSON - .apply(pd.Series) As an alternative solution we can use `.apply(pd.Series)`. The result is the same: ```python df['col_json'].apply(pd.Series) ``` this will result into new DataFrame with values stored in the JSON: | | x | y | xy | | - | - | - | -- | | 0 | 1 | 0 | 1 | | 1 | 0 | 1 | 1 | | 2 | 1 | 1 | 1 | × **Pro Tip** Method **pd.json\_normalize** is faster than .apply(pd.Series) \- tested by \`%%timeit\` Compare performance of `json_normalize` and `.apply(pd.Series)` : - `json_normalize` \- 361 µs ± 2.99 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) - `.apply(pd.Series)` \- 1.29 ms ± 21.7 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) ## 3: Parse JSON - json.loads + ast.literal\_eval Methods `ast.literal_eval` and `json.loads` help us to parse JSON data. They can help us when we need to read and parse JSON stored as string. Let's see them how they work and what is the key difference: ```python json.loads('{"x": 1}') ``` result: ``` {'x': 1} ``` ```python ast.literal_eval('{"x": 1}') ``` we get the same result: ``` {'x': 1} ``` × **Pro Tip** Method **json.loads** can work only with valid JSONs while ast.literal_eval can load wider variety of JSON like syntax Which means that: ```python json.loads("{'x': 1}") ``` will error with: > JSONDecodeError: Expecting property name enclosed in double quotes: line 1 column 2 (char 1) But we can convert **non standard JSON data by `ast.literal_eval`**: ```python ast.literal_eval("{'x': 1}") ``` ## 4: Flattening JSON stored as string What if we like to **normalize JSON which is stored as string in Pandas column**. Using previous steps will not help. It will result in a single column named `0`. In order to **convert dict or JSON stored as a string to multiple columns** we can use combination of: - `pd.json_normalize` - `ast.literal_eval` or `json.loads` (for strict JSON format) ```python import ast df['col_str_dict'].apply(ast.literal_eval) ``` This will result into: | | x | y | xy | | - | - | - | -- | | 0 | 1 | 0 | 1 | | 1 | 0 | 1 | 1 | | 2 | 1 | 1 | 1 | ## 5: Flattening nested JSON - pd.json\_normalize Suppose we have data as follow: ```json data = [ { 'pos': 'east', 'population': 55, 'continent': 'Africa', 'info': { 'people': { 'environmentalist': 'Wangari Maathai', 'journalist': 'Ngugi wa Thiong''o' } }, 'city': [ { 'name': 'Nairobi', 'capital': True, 'population': { 'total': 4.4, 'rank': 11 } }, { 'name': 'tigerfish', 'capital': False, 'population': { 'total': 1.2, 'rank': 12 } }, ] }, { 'pos': 'central', 'population': 218, 'continent': 'Africa', 'people': { 'famous': { 'writer': 'Ken Saro-Wiwa', 'artist': 'Tiwa Savage' } }, 'city': [ { 'name': 'Lagos', 'capital': True }, { 'name': 'Kano', 'capital': False }, ] } ] ``` The table represents the data stored as a DataFrame: | | pos | population | continent | info | city | people | | - | ------- | ---------- | --------- | ------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------- | | 0 | east | 55 | Africa | {'people': {'environmentalist': 'Wangari Maathai', 'journalist': 'Ngugi wa Thiongo'}} | \[{'name': 'Nairobi', 'capital': True, 'population': {'total': 4.4, 'rank': 11}}, {'name': 'tigerfish', 'capital': False, 'population': {'total': 1.2, 'rank': 12}}\] | NaN | | 1 | central | 218 | Africa | NaN | \[{'name': 'Lagos', 'capital': True}, {'name': 'Kano', 'capital': False}\] | {'famous': {'writer': 'Ken Saro-Wiwa', 'artist': 'Tiwa Savage'}} | If we need to specify the path and the metadata we can do: ```python pd.json_normalize( data, record_path =['city'], meta=['pos', 'continent', 'population'], errors='ignore' ) ``` which will result into normalized form: | | name | capital | population.total | population.rank | pos | continent | population | | - | --------- | ------- | ---------------- | --------------- | ------- | --------- | ---------- | | 0 | Nairobi | True | 4.4 | 11.0 | east | Africa | 55 | | 1 | tigerfish | False | 1.2 | 12.0 | east | Africa | 55 | | 2 | Lagos | True | NaN | NaN | central | Africa | 218 | | 3 | Kano | False | NaN | NaN | central | Africa | 218 | ## 6: Flattening a JSON with multiple levels There is a Python library accessible at: [flatten-json](https://pypi.org/project/flatten-json/?ref=datascientyst.com). It can be installed by: ```python pip install flatten-json ``` The library is described as: > Flattens JSON objects in Python. flatten\_json flattens the hierarchy in your object which can be useful if you want to force your objects into a table. For this step we are going to create additional DataFrame: ```python data = {'col_json': {0: {"x": 1.0, "y": 0.0}, 1: {"x": 0.0, "y": 1.0}, 2: {"x": 1.0, "y": 1.0}} } df2 = pd.DataFrame(data) ``` To demonstrate how we can **flatten JSON objects in Python and Pandas:** ```python from flatten_json import flatten df2['col_json'].apply(flatten).apply(pd.Series) ``` the result: | | x | y | | - | --- | --- | | 0 | 1.0 | 0.0 | | 1 | 0.0 | 1.0 | | 2 | 1.0 | 1.0 | ## 7: Parse local/remote JSON file If JSON data is stored as a file - locally or remotely we can normalize it with few additional lines: ### Normalize local JSON file in Pandas The code below will load and normalize local file in the same folder as the script: ```python import json with open('data.json','r') as f: data = json.loads(f.read()) pd.json_normalize(data) ``` ### Normalize remote JSON file in Python In order to load and normalize JSON data from a remote file we can use the following code: ```python import requests URL = 'https://raw.githubusercontent.com/softhints/Pandas-Tutorials/master/data/sample.json' data = json.loads(requests.get(URL).text) pd.json_normalize(data) ``` ## 8: Concatenate expanded data to DataFrame If you need to **concatenate the normalized or flattened data to the original DataFrame** we can use method `concat`: ```python import json pd.concat([df, df['col_str_dict_2'].apply(json.loads).apply(pd.Series)], axis=1) ``` and apply the operations on columns: | | col\_json | col\_str\_dict | col\_str\_dict\_list | x | | - | ------------------------- | ------------------------ | -------------------- | -------- | | 0 | {'x': 1, 'y': 0, 'xy': 1} | {'x': 1, 'y': 0, 'xy':1} | {"x": \[1,1\]} | \[1, 1\] | | 1 | {'x': 0, 'y': 1, 'xy': 1} | {'x': 0, 'y': 1, 'xy':1} | {"x": \[0,1\]} | \[0, 1\] | | 2 | {'x': 1, 'y': 1, 'xy': 1} | {'x': 1, 'y': 1, 'xy':1} | {"x": \[1,0\]} | \[1, 0\] | ## 9\. Multiple levels, custom parsing Sometimes we might have multiple levels and errors. We can bypass the errors from json conversion by using custom method: ```python import ast import pandas as pd import json def parse_eval(value): try: return ast.literal_eval(value) except (ValueError, SyntaxError): return value pd.DataFrame(json.loads("[" + df1['items'].str.replace("\'", "").replace(r"\\\"", "", regex=False).apply(parse_eval).apply(pd.Series)[0].iloc[0] + "]")).title ``` In the code above: - we load content of column items - then fix the extra characters like `'` and `"` - parse the json - expand nested values - get the content - load the string as json (adding parathensis) - create dataframe - finally extra field title only Often you need to apply custom steps depending on your case. ## More examples How to flatten nested JSON in Pandas: ```python import ast df['col_str_dict_2'].apply(ast.literal_eval).apply(pd.Series) ``` Normalize semi-structured JSON data into a flat table: ```python import ast df['col_str_dict_2'].apply(pd.json_normalize) ``` Fix multiple errors with the input JSON data and format: - replace quotes - replace unicode characters - load as JSON - convert to columns ```python df_spm['result'].str.replace('\\"', '', regex=False).replace("\\\'", '', regex=True)\ .replace(u'‎\u200e', '', regex=True).str.encode('ascii','ignore')\ .apply(json.loads).apply(pd.Series) ``` ## TypeError: the JSON object must be str, bytes or bytearray, not 'float' Sometimes we might get errors like: - `TypeError: the JSON object must be str, bytes or bytearray, not 'float'` - `TypeError: the JSON object must be str, bytes or bytearray, not dict` Different reasons can cause similar errors. This error will be raised if we try to apply `json.loads` to a JSON data: ```python df2['col_json'].apply(json.loads) ``` To avoid such errors we might convert the column to string or parse it by library `flatten_json`: ```python from flatten_json import flatten df2['col_json'].apply(flatten).apply(pd.Series) ``` ## TypeError: string indices must be integers This error may be the result of misuse of the method: `pd.json_normalize()`. Be sure to pass JSON data. If we are passing DataFrame then we need to convert it to proper JSON by: ```python data.to_dict() ``` ### ValueError: malformed node or string on line 1 If you face Pandas error: > ValueError: malformed node or string on line 1: while using `.apply(ast.literal_eval)`. Then you can try: ```python df_spm['result'].str.replace('\\"', '', regex=False).replace("\\\'", '', regex=True)\ .replace(u'‎\u200e', '', regex=True).str.encode('ascii','ignore')\ .apply(json.loads).apply(pd.Series) ``` ## Conclusion In this article we covered multiple ways to convert JSON data or columns containing JSON data to multiple columns. We discussed different problems and solutions of most typical problems. It mentioned performance benefits and working with multiple level JSON data. ### 108-cheat-sheet URL: https://datascientyst.com/108-cheat-sheet/ Last updated: 2026-08-31T23:11:01.000Z Cheat Sheet ### How to Convert String to DateTime in Pandas URL: https://datascientyst.com/convert-string-to-datetime-pandas/ Last updated: 2022-10-05T08:33:03.000Z To **convert string column to DateTime in Pandas and Python** we can use: **(1) method: `pd.to_datetime()`** ```python pd.to_datetime(df['date']) ``` **(2) method .astype('datetime64\[ns\]')** ```python df['timestamp'].astype('datetime64[ns]') ``` Let's check the most popular cases of conversion of string to dates in Pandas like: - custom date and time patterns - infer date pattern - dates different language locales - different date formats ![](https://datascientyst.com/content/images/2022/06/convert-string-to-datetime-pandas.png) ## Setup Suppose we have DataFrame with Unix timestamp column as follows: ```python dict = {'date': {0: '28-01-2022 5:25:00 PM', 1: '27-02-2022 6:25:00 PM', 2: '30-03-2022 7:25:00 PM', 3: '29-04-2022 8:25:00 PM', 4: '31-05-2022 9:25:00 PM'}, 'date_short': {0: 'Jan-2022', 1: 'Feb-2022', 2: 'Mar-2022', 3: 'Apr-2022', 4: 'May-2022'}} df = pd.DataFrame(dict) ``` So data will look like: | | date | date\_short | | - | --------------------- | ----------- | | 0 | 28-01-2022 5:25:00 PM | Jan-2022 | | 1 | 27-02-2022 6:25:00 PM | Feb-2022 | | 2 | 30-03-2022 7:25:00 PM | Mar-2022 | | 3 | 29-04-2022 8:25:00 PM | Apr-2022 | | 4 | 31-05-2022 9:25:00 PM | May-2022 | ## Step 1: Convert string to date with pd.to\_datetime() The first and the most common example is to convert a time pattern to a datetime in Pandas. To do so we can use method `pd.to_datetime()` which will recognize the correct date in most cases: ```python pd.to_datetime(df['date']) ``` The result is the correct datetime values: ``` 0 2022-01-28 17:25:00 1 2022-02-27 18:25:00 2 2022-03-30 19:25:00 3 2022-04-29 20:25:00 4 2022-05-31 21:25:00 Name: date, dtype: datetime64[ns] ``` ## Step 2: Convert time or date pattern "%d/%m/%Y" to date The method `to_datetime` has different parameters which can be found on: [pandas.to\_datetime](https://pandas.pydata.org/docs/reference/api/pandas.to%5Fdatetime.html?ref=datascientyst.com). To give a date format we can use parameter `format`: ```python pd.to_datetime('20220701', format='%Y%m%d', errors='ignore') ``` Once more example: ```python pd.to_datetime(df['date'] , format='%Y%m%d HH:MM:SS', errors='ignore') ``` Note: If we use wrong format we will get an error: > ValueError: time data '28-01-2022 5:25:00 PM' does not match format '%Y%m%d HH:MM:SS' (match) In order to solve it we can use `errors='ignore'`. To understand how to analyze Pandas date errors you can check this article: [OutOfBoundsDatetime: Out of bounds nanosecond timestamp - Pandas and pd.to\_datetime ](https://datascientyst.com/outofboundsdatetime-out-of-bounds-nanosecond-timestamp-pandas-pd-to%5Fdatetime/) To find more Pandas errors related to dates please check: [Pandas Most Typical Errors and Solutions for Beginners](https://datascientyst.com/pandas-most-typical-errors-and-solutions/) ## Step 3: Check if string is a date in Pandas If we need to **check if a given column contain dates** (even if there are extra characters or words) we can build method like: ```python from dateutil.parser import parse def is_date(string, fuzzy=False): try: parse(string, fuzzy=fuzzy) return True except ValueError: return False ``` Then we can use it as: ```python df['date'].apply(is_date) ``` For fuzzy matching - meaning that: ``` today is 2019-03-27 ``` will return True - we need to call it as: ```python df['date_short'].apply(is_date, fuzzy=True) ``` ## Step 4: Infer date format from string Python and Pandas has several option if we need to **infer the date or time pattern from a string**. ### \_guess\_datetime\_format\_for\_array The first option is by using `_guess_datetime_format_for_array`: ```python import numpy as np from pandas.core.tools.datetimes import _guess_datetime_format_for_array array = np.array(['2022-06-01T00:10:45.300000']) _guess_datetime_format_for_array(array) ``` which will result into: ``` '%Y-%m-%dT%H:%M:%S.%f' ``` This option has some limitations and might return `None` for valid dates. For Pandas column we can use: ```python import numpy as np from pandas.core.tools.datetimes import _guess_datetime_format_for_array array = np.array(df["Date"].to_list()) _guess_datetime_format_for_array(array) ``` ### hi-dateinfer - Before python 3.8 We can use library: [hi-dateinfer](https://pypi.org/project/hi-dateinfer/?ref=datascientyst.com) which can be installed by: ```bash pip install hi-dateinfer ``` Now we can infer date or time format for Pandas column as follows: ```python import hidateinfer as dateinfer df['Date'].apply(dateinfer.infer) ``` Which would give us something like: ``` '%a %b %d %H:%M:%S %Z %Y' ``` ### py-dateinfer - Before python 3.8 Another option is to use Python library: `py-dateinfer` which can be installed by: ```bash pip install py-dateinfer ``` To use it we can do: ```python import dateinfer dateinfer.infer(['Mon Jan 13 09:52:52 MST 2014', 'Tue Jan 21 15:30:00 EST 2014']) ``` the result would be: ``` '%a %b %d %H:%M:%S %Z %Y' ``` ## Step 5: Generic parsing of dates 200 locales What if we need to **parse dates in different languages like**: - French - Thai - Russian - Spanish In this case we can use the Python library called `dateparser`. It can be installed by: ```python pip install dateparser ``` To parse different locales with dateparser and Pandas: ```python df['date'].apply(dateparser.parse) ``` We can also use settings: ```python df['dates'].apply(dateparser.parse, settings={'DATE_ORDER': 'DMY'}) ``` ## Step 6: Working with mixed datetime formats Finally lets cover the case of **multiple date formats in a single column in Pandas.** In that case we can build a custom function to detect a parse the correct format like: ```python def date_parser(date): if '/' in date: return pd.to_datetime(date, format = '%Y/%m/%d') elif '-' in date: return pd.to_datetime(date, format = '%d-%m-%Y') else: return pd.to_datetime(date, format = '%d.%m.%Y') ``` Or we can parse different format separately and them merge the results: ```python date_1 = pd.to_datetime(df['date'], errors='coerce', format='%Y/%m/%d') date_2 = pd.to_datetime(df['date'], errors='coerce', format='%d-%m-%Y') df['date'] = date_1.fillna(date_2) ``` ## Conclusion In this article we covered conversion of string to date in Pandas. We covered multiple edge cases like locales, formats and errors. Now we know how to infer date format from a string and how to parse multiple formats in a single Pandas column. ### How to Convert Unix Time to Date in Pandas URL: https://datascientyst.com/convert-unix-time-to-date-pandas/ Last updated: 2022-06-23T04:07:37.000Z To **convert Unix timestamp to readable date in Pandas** we can use method: `pd.to_datetime` ```python df['date'] = pd.to_datetime(df['date'],unit='s') ``` So this will convert: ``` [1655822072.437469, 1655815574.333629, 1655797456.516109] ``` to datetime in Pandas: ``` DatetimeIndex(['2022-06-21 14:34:32.437469006', '2022-06-21 12:46:14.333628893', '2022-06-21 07:44:16.516108990'], dtype='datetime64[ns]', freq=None) ``` Let's cover all the steps in to practical example - **converting Unix timestamp to any date format (including dd/mm/yyyy)**. ## Setup Suppose we have DataFrame with Unix timestamp column as follows: ```python dict = {'ts': {0: 1655822072.437469, 1: 1655815574.333629, 2: 1655797456.516109, 3: 1655743965.358579, 4: 1655712623.707739}, 'reply_count': {0: 2.0, 1: 3.0, 2: 3.0, 3: 2.0, 4: None}} pd.DataFrame(dict) ``` So data will look like: | | ts | reply\_count | | - | ----------------- | ------------ | | 0 | 1655822072.437469 | 2.0 | | 1 | 1655815574.333629 | 3.0 | | 2 | 1655797456.516109 | 3.0 | | 3 | 1655743965.358579 | 2.0 | | 4 | 1655712623.707739 | NaN | ![](https://datascientyst.com/content/images/2022/06/convert-unix-time-to-date-pandas.png) ## Step 1: Convert Unix time column to datetime The first step is to convert the Unix timestamp to Pandas datetime by: ```python df['date'] = pd.to_datetime(df['ts'], unit='s') ``` The important part of the conversion is `unit='s'` which stands for seconds. There other options like: - `ns` \- nanoseconds - `ms` \- milliseconds Default value is None and all available options can be found here: [pandas.Timestamp](https://pandas.pydata.org/docs/reference/api/pandas.Timestamp.html?ref=datascientyst.com) × **Pro Tip 1** Sometimes the Unix time can be stored as a string - so conversion to integer may be needed: .astype(int) ```python df['ts'] = df['ts'].astype(int) ``` ## Step 2: Convert Unix time to readable date The second step is to convert Pandas datetime to a readable date. This is possible by using `dt` attribute: ```python df['date'].dt.date ``` The output will be the date component of the original Unit time: ``` 0 2022-06-21 1 2022-06-21 2 2022-06-21 3 2022-06-20 4 2022-06-20 ``` ## Step 3: Convert Unix time to readable time To convert the Unix time to a well formatted time string we can use again the `dt` attribute: ```python df['date'].dt.time ``` will give us: ``` 0 14:34:32.437468 1 12:46:14.333628 2 07:44:16.516109 3 16:52:45.358578 4 08:10:23.707739 ``` ## Step 4: Convert Unix time to custom date or time format Suppose we would like to get different time pattern like: - `dd/mm/yy` - `HH:MM` etc This is possible by using method `.dt.strftime()`: ```python df['date'].dt.strftime('%m/%Y') ``` which will result into: ``` 0 06/2022 1 06/2022 2 06/2022 3 06/2022 4 06/2022 ``` To find more examples you can consult with: [How to Extract Month and Year from DateTime column in Pandas ](https://datascientyst.com/extract-month-and-year-datetime-column-in-pandas/) ## Step 5: Use datetime.datetime.utcfromtimestamp Alternative solution is to use `datetime.datetime.utcfromtimestamp` to convert Unix timestamp to date in Pandas. To use method like `datetime.utcfromtimestamp` we will need to apply it to the Unix column: ```python from datetime import datetime df["ts"].apply(lambda x: datetime.utcfromtimestamp(x).strftime('%Y-%m-%dT%H:%M:%SZ')) ``` In this way we can specify the format like: - `%Y-%m-%dT%H:%M:%SZ` - `%d-%m-%Y %H:%M:%S` ## Conclusion In this article, we saw multiple ways to convert timestamp columns to datetime. We also covered multiple date and time formats, plus possible problems. ### Pandas vs SQL Cheat Sheet URL: https://datascientyst.com/pandas-vs-sql-cheat-sheet/ Last updated: 2023-04-01T06:57:29.000Z With this **SQL & Pandas cheat sheet**, we'll have a valuable reference guide for Pandas and SQL. We can convert or run SQL code in Pandas or vice versa. Consider it as **Pandas cheat sheet for people who know SQL**. The **cheat sheet covers basic querying tables, filtering data, aggregating data, modifying and advanced operations.** It includes the most popular operations which are used on a daily basis with SQL or Pandas. Keep reading until the end for bonus tips and conversion between Pandas and SQL. ## Select Selecting data with Pandas vs SQL ![](https://datascientyst.com/content/images/2022/06/sql_pandas_cheat_sheet_select_o.png) `df` SELECT \* FROM tab; query data from all columns `df[['col_1', 'col_2']]` SELECT col\_1, col\_2 FROM tab; select subset of columns and get all rows `df.assign(new_col = df['col_1'] / df['col_2'])` SELECT \*, col\_1/col\_2 as new\_col FROM tab; Aliases - add a calculated column based on other columns `df[['col 1']]` SELECT \`col 1\` FROM tab; Column name with space - \`col 1\` `df.sort_values(by='col_1', ascending=True)` SELECT \* FROM tab ORDER BY col\_1 ASC; Sort values by col\_1 in ascending or descending (DESC) order ## Where Filtering in SQL vs Pandas ![](https://datascientyst.com/content/images/2022/06/sql_pandas_cheat_sheet_where_o.png) `df[df['col_1'] == '11']` SELECT \* FROM tab WHERE col\_1 = '11'; Filtering on single condition `df[(df['col_1'] == 11) & (df['col_2'] > 5)]` SELECT \* FROM tab WHERE col\_1 = 11 AND col\_2 > 5; Filtering on multiple conditions `df[df['col_2'].isna()]` SELECT \* FROM tab WHERE col\_2 IS NULL; NULL checking is done using the isna() `df[df['col_1'].notna()]` SELECT \* FROM tab WHERE col\_1 IS NOT NULL; NOT NULL checking is done using the notna() `df[df.col_1 > df.col_2]` SELECT \* FROM tab WHERE col\_1 > col\_2; Where clause with 2 SQL columns `df.query('col_1 == col_2')` SELECT \* FROM tab WHERE col\_1 = col\_2; Filter with Pandas Query `` df.query('col_1 == `col 2`') `` SELECT \* FROM tab WHERE col\_1 = \`col 2\`; Column name with space - \`col 2\` in where clause ## Like, and, or Operators(Text, Logical) in Pandas vs SQL ![](https://datascientyst.com/content/images/2022/06/sql_pandas_cheat_sheet_like_or_and_o.png) `df[df['col_1'].str.contains('i', na=False)]` WHERE col\_1 LIKE '%i%' Finds values which contain \`i\`. Column need to be string - .astype(str) `df[df['col_1'].str.contains('sh|rd', regex=True, na=True)]` WHERE col\_1 LIKE '%sh%' OR col\_1 LIKE '%rd%' Finds any values which contain \`sh\` or \`rd\` `df[df['col_1'].str.startswith('h', na=False)]` WHERE col\_1 LIKE 'h%' Finds any values that start with \`h\` `df[df['col_1'].astype(str).str.endswith('k', na=False)]` WHERE col\_1 LIKE '%k' Finds any values that ends with \`k\` `(df['col_1'] == '11') & (df['col_2'] > 5)` WHERE col\_1 = '11' AND col\_2 > 5; AND = SQL - \`and\`, Pandas - \`&\` `(df['col_1'] == '11') | (df['col_2'] > 5)` WHERE col\_1 = '11' OR col\_2 > 5; OR = SQL - \`or\`, Pandas - \`|\` `df[df['col_1'].isin([1,2,3])]` SELECT \* FROM tab WHERE col\_1 in (1,2,3); IN operator - find values from list of values `df[df['col_1'].between(1, 5)]` SELECT \* FROM tab WHERE col\_1 BETWEEN 1 AND 5; BETWEEN operator - find values in a range ## Group by Group by operations in SQL vs Pandas ![](https://datascientyst.com/content/images/2022/06/sql_pandas_cheat_sheet_group_by_o.png) `df.groupby('col_1').size() df.groupby('col_1')['col_2'].count() df.col_1.value_counts()` SELECT col\_1, count(\*) FROM tab GROUP BY col\_1; count records in each group( 3 versions in Pandas ) `import numpy as np df.groupby('col_1').agg({'col_2': np.mean, 'col_1': np.size})` SELECT col\_1, AVG(col\_2), COUNT(\*) FROM tab GROUP BY col\_1; Apply multiple statistical functions `df.groupby(['col_1', 'col_2']).agg({'col_3': [np.size, np.mean]})` SELECT col\_1, col\_2, COUNT(\*), AVG(col\_3) FROM tab GROUP BY col\_1, col\_2; Grouping by multiple columns, multiple functions `# group by g = df.groupby('col_1') # having count(*) > 10 g.filter(lambda x: len(x) > 10)['col_1']` SELECT col\_1, count(\*) FROM tab GROUP BY col\_1 HAVING count(\*) > 10; HAVING - Group by column and filtering contidion on the groups `df.groupby('col_1').col_2.nunique()` SELECT count(distinct col\_2) FROM tab GROUP BY col\_1; count(distinct) - count unique elements in group ## Join Join in SQL and Pandas ![](https://datascientyst.com/content/images/2022/06/sql_pandas_cheat_sheet_join_o.png) `pd.merge(df1, df2, on='key')` SELECT \* FROM t1 INNER JOIN t2 ON t1.key = t2.key; Inner join of 2 table/dataframes(1) `pd.merge(df1, df2, on='key', how='left')` SELECT \* FROM t1 LEFT OUTER JOIN t2 ON t1.key = t2.key; Left Outer join `pd.merge(df1, df2, on='key', how='right')` SELECT \* FROM t1 RIGHT OUTER JOIN t2 ON t1.key = t2.key; Right Join `pd.merge(df1, df2, on='key', how='outer')` SELECT \* FROM t1 FULL OUTER JOIN t2 ON t1.key = t2.key; Full Join (not working on MySQL) `pd.merge(df1, df2, left_on= ['col_1', 'col_2'], right_on= ['col_3', 'col_4'], how = 'right')` SELECT \* FROM t1 INNER JOIN t2 ON t1.col\_1 = t2.col\_3 AND t1.col\_2 = t2.col\_4; Join on columns with different names `m = pd.merge(df1, df2, how='left', on=['col_1', 'col_2']) pd.merge(m, df3[['col_1', 'col_2', 'col_3']], how='left', on=['col_1', 'col_3'])` SELECT t1.col\_a, t2.col\_b, t3.col\_c FROM t1 LEFT OUTER JOIN t2 ON t1.col\_1 = t2.col\_1 AND t1.col\_2 = t2.col\_2 LEFT OUTER JOIN t3 ON t1.col\_1 = t3.col\_1 AND t1.col\_3 = t3.col\_3 join multiple dataframes on multiple columns ## Union Union in SQL and Pandas ![](https://datascientyst.com/content/images/2022/06/sql_pandas_cheat_sheet_union_o.png) `pd.concat([df1, df2])` SELECT \* FROM df1 UNION ALL SELECT \* FROM df2; Unioun All(columns must have same number of columns) `cols= ['col_1', 'col_2'] pd.concat([df1[cols], df2[cols]])` SELECT col\_1, col\_2 FROM t1 UNION ALL SELECT col\_1, col\_2 FROM t2; Unioun All `pd.concat([df1, df2]).drop_duplicates()` SELECT col\_1, col\_2 FROM t1 UNION SELECT col\_1, col\_2 FROM t1; Unioun All ( remove duplicate rows) ## Limit Limit in SQL and Pandas ![](https://datascientyst.com/content/images/2022/06/sql_pandas_cheat_sheet_limit_o.png) `df.head(10) ` SELECT \* FROM tab LIMIT 10; Get top rows `df.tail(10)` SELECT \* FROM tab ORDER BY id DESC LIMIT 10; Get last N rows `lim = 2 offset = 5 df.sort_values('col_1', ascending=False).iloc[offset:lim+offset]` SELECT \* FROM tab ORDER BY col\_1 DESC LIMIT 2 OFFSET 5; return only top 2 records, start on record 6 (OFFSET 5) ## Update Update in SQL vs Pandas ![](https://datascientyst.com/content/images/2022/06/sql_pandas_cheat_sheet_update_o.png) `df.loc[df['col_1'] < 2, 'col_1'] *= 2` UPDATE tab SET col\_1 = col\_1\*2 WHERE col\_1 < 2; Update all rows for 1 column with condition `df1['name'] = np.where(df2['id']==1,df2['name'],df1['name'])` UPDATE t1, ( SELECT \* FROM t2 WHERE id = 1 ) AS temp SET t1.name = temp.name WHERE t1.id = 1; Update based on select from another table / dataframe ## Delete Delete in SQL vs Pandas ![](https://datascientyst.com/content/images/2022/06/sql_pandas_cheat_sheet_delete_o.png) `df = df.loc[df['col_1'] > 9]` DELETE FROM tab WHERE col\_1 > 9; Delete rows with condition `df1.drop(df1[(df1.id.isin(df2.id) & (df1.id==1))].index)` DELETE t1 FROM df1 as t1 JOIN df2 as t2 ON t1.id = t2.id WHERE t2.id = 1; Delete rows with condition based on another table / dataframe ## Insert Insert in SQL vs Pandas ![](https://datascientyst.com/content/images/2022/06/sql_pandas_cheat_sheet_insert_o.png) `data = {'col_1': 1, 'col_2': '11'} df = df.append(data, ignore_index = True)` INSERT INTO tab(col\_1, col\_2) VALUES (1, '11'); Add new data/rows to a table/dataframe Now we can explore the conversion between Pandas and SQL. The SQL syntax is based and tested on MySQL 8. Additional resources: - [Pandas Cheat Sheet for Data Science](https://datascientyst.com/pandas-cheat-sheet-for-data-science/) - [Pandas Comparison with SQL](https://pandas.pydata.org/docs/getting%5Fstarted/comparison/comparison%5Fwith%5Fsql.html?ref=datascientyst.com) ## Pandas equivalent of SQL Table below shows the most used equivalents between SQL and Pandas: | Category | Pandas | SQL | | -------------------- | ---------------------------------------- | ----------------------------- | | Data Structure | DataFrames, Series | Tables, Row, Column | | Querying | .loc , .iloc , boolean indexing | SELECT, WHERE | | Filtering | .query() , boolean indexing | WHERE | | Sorting | .sort\_values() , .sort\_index() | ORDER BY | | Grouping | .groupby() , .agg() | GROUP BY, AGGREGATE | | Joins | .merge() , .join(), .concat() | JOIN, UNION | | Aggregations | .agg() , .apply() | AGGREGATE | | Data Transformation | .apply() , .map() , .replace() | replace(column, 'old', 'new') | | Count Unique | .nunique(), .agg(\['count', 'nunique'\]) | count(distinct) | | Data Cleaning | .dropna() , .fillna() | IS NULL; IS NOT NULL; | | Data Type Conversion | .astype() | CAST | ## Differences between SQL and Pandas Usually there are multiple ways to achieve something in Pandas or SQL. Pandas offers a bigger variety of options. Often the solution depends on the performance. There are some key differences when you work with SQL or Pandas: ### Join in SQL vs Pandas (1) In Pandas if both key columns contain rows with `NULL` value, those rows will be matched against each other. This is not the case for SQL join behavior and can lead to unexpected results in Pandas ### Copies vs. in place operations Usually in Pandas operations return copies of the Series/DataFrame. To change this behavior we can use: ```python df.sort_values("col_1", inplace=True) ``` Create a new DataFrame: ```python sorted_df = df.sort_values("col1") ``` or update the original one: ```python df = df.sort_values("col1") ``` ## Pandas to SQL To **convert or export Pandas DataFrame to SQL** we can use method: `to_sql()`: There are several important parameters which need to be used for this method: - `name` \- table name in the database - `con` \- DB connection. sqlalchemy.engine.(Engine or Connection) or sqlite3.Connection - `if_exists` \- behavior if the table exists in the DB - `index` \- convert DataFrame index to a table column - `chunksize` \- number of rows in each batch to be written at a time A very basic example is given below: ```python from sqlalchemy import create_engine engine = create_engine('sqlite://', echo=False) df = pd.DataFrame({'name' : ['User 1', 'User 2', 'User 3']}) df.to_sql('users', con=engine) ``` We can check results by using SQL query like: ```python engine.execute("SELECT * FROM users").fetchall() ``` result: ``` [(0, 'User 1'), (1, 'User 2'), (2, 'User 3')] ``` The full documentation is available on this link: [to\_sql()](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to%5Fsql.html?ref=datascientyst.com) ## Pandas from SQL Pandas offers method `read_sql()` to **get data from SQL or DB to a DataFrame or Series**. The method can be used to read SQL connection and fetch data: ```python pd.read_sql('SELECT col_1, col_2 FROM tab', conn) ``` where `conn` is SQLAlchemy connectable, str, or sqlite3 connection. To find more you can check: [pandas.read\_sql()](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fsql.html?ref=datascientyst.com) ## Pandas and SQL with SQLAlchemy and PyMySQL Alternatively we can **convert SQL table to Pandas DataFrame with SQLAlchemy and PyMySQL**. To do so check: [How to Convert MySQL Table to Pandas DataFrame / Python Dictionary](https://softhints.com/convert-mysql-table-pandas-dataframe-python-dictionary/?ref=datascientyst.com) It includes many different step by step examples and options. You can find how to connect Python to SQL using libraries - SQLAlchemy and PyMySQL. ## Run SQL code in Pandas To run SQL syntax in Python and Pandas we need to install Python package - `pandasql` by: ```python pip install -U pandasql ``` The package allows us to **query Pandas DataFrames using SQL syntax**. It uses SQLite syntax. More information can be found here - [pandasql](https://pypi.org/project/pandasql/?ref=datascientyst.com). Below we can see a basic example of running SQL syntax to Pandas DataFrame : ```python from pandasql import sqldf import pandas as pd q = "SELECT * FROM df LIMIT 5" sqldf(q, globals()) ``` Method `sqldf()` requires 2 parameters: - the SQL query - `globals()` or `locals()` function ## Generate SQL statements in Pandas ### SQL create table To **generate SQL statements from Pandas DataFrame** We can use the method `pd.io.sql.get_schema()`. So the get the create SQL statement for a given DataFrame we can use: ```python pd.io.sql.get_schema(df.reset_index(), 'tab') ``` where 'tab' is the name of the DB table. ### SQL insert table Generating SQL insert statements from a DataFrame can be achieved by: ```python sql_insert = [] for index, row in df.iterrows(): sql_insert.append('INSERT INTO `tab` ('+ str(', '.join(df.columns))+ ') VALUES '+ str(tuple(row.values))) ``` Where: - `df` is the name of the source DataFrame - `tab` is the target table We are iterating over all rows of the DataFrame and generating insert statements. ## Conclusion We covered basic and advanced operations with Pandas and SQL. You can use this to **quickly transfer SQL to Pandas and the reverse.** If you like content like this - make sure to subscribe to get the latest updates and resources. This post is only a part from a serires related to cheat sheets related to data science. ![](https://datascientyst.com/content/images/2022/06/pandas_vs_sql_cheatsheet-2.png) ### How to Swap Levels of MultiIndex in Pandas URL: https://datascientyst.com/swap-levels-multiindex-pandas/ Last updated: 2022-06-08T06:46:51.000Z In this short guide, I'll show you how to **swap levels of MultiIndex in Pandas DataFrame**. You can also find how to reorder MultiIndex from left to right(and reverse), swap multiple levels and typical errors. So at the end we will get different order of the MultiIndex levels from: ``` MultiIndex([('11', '21', '31'), ('11', '22', '32'), ('12', '21', '33'), ('12', '22', '34')], ) ``` to: ``` MultiIndex([('11', '31', '21'), ('11', '32', '22'), ('12', '33', '21'), ('12', '34', '22')], ) ``` Notice that the inner levels changed their positions. This is done by method: ```python df.swaplevel() ``` Let's cover the MultiIndex reorder in more detail. ## Setup To start let's create a simple DataFrame with MultiIndex: ```python import pandas as pd df = pd.DataFrame( {"Grade": ["A", "B", "A", "C"]}, index=[ ["11", "11", "12", "12"], ["21", "22", "21", "22"], ["31", "32", "33", "34"] ] ) ``` The resulted DataFrame is: | | | | Grade | | -- | -- | -- | ----- | | 11 | 21 | 31 | A | | 22 | 32 | B | | | 12 | 21 | 33 | A | | 22 | 34 | C | | We can get the index by: ```python df.index ``` result is a MultiIndex: ``` MultiIndex([('11', '21', '31'), ('11', '22', '32'), ('12', '21', '33'), ('12', '22', '34')], ) ``` ## Step 1: Method swaplevel() We can use the method `swaplevel()` to swap levels of DataFrame with MultiIndex. The method's documentation is available from: [DataFrame.swaplevel](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.swaplevel.html?ref=datascientyst.com). The method signature is: ```python DataFrame.swaplevel(i=- 2, j=- 1, axis=0) ``` Where the parameters are: - `i, j` \- Levels of the indices to be swapped (int or str) - `axis` \- 0 or `index`, 1 or `columns` \- column-wise We covered the result of the default behavior above. Here we will explain it. Default behavior ```python df.swaplevel() ``` is equivalent of: ```python df.swaplevel(i=-2, j=-1, axis=0) ``` So it will swap the two innermost levels row-wise. ![](https://datascientyst.com/content/images/2022/06/swap-levels-multiindex-pandas.png) ## Step 2: Swap Levels of MultiIndex - column-wise **Swapping levels of MultiIndex column-wise** is possible by using parameter `axis=1`. ```python df.swaplevel(axis=1) ``` ## Step 3: Swap Outer Levels of MultiIndex Levels of the MultiIndex: ``` MultiIndex([('11', '21', '31'), ('11', '22', '32'), ('12', '21', '33'), ('12', '22', '34')], ) ``` Are numbered as: - 0 - 1 - 2 To get level 0 we can do: ```python df.index.get_level_values(0) ``` The result is index: ``` Index(['11', '11', '12', '12'], dtype='object') ``` So to **swap the levels from the left to the right we can use positive numbers**: ```python df.swaplevel(i=0, j=1) ``` From right to the left we can pass negative indices: ```python df.swaplevel(i=-1, j=-2) ``` × **Pro Tip 1** df.swaplevel(i=-1, j=-2) is equivalent to df.swaplevel(i=-2, j=-1) ## Step 4: reorder\_levels() - swap multiple levels Method `reorder_levels()` can be used to **swap multiple levels of MultiIndex** at once. It requires the level order as parameter: ```python df.reorder_levels([1,2,0]).index ``` result: ``` MultiIndex([('21', '31', '11'), ('22', '32', '11'), ('21', '33', '12'), ('22', '34', '12')], ) ``` You can find more about this method here: [DataFrame.reorder\_levels](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.reorder%5Flevels.html?ref=datascientyst.com) ## Step 5: TypeError: Can only swap levels on a hierarchical axis. If you get error like: > TypeError: Can only swap levels on a hierarchical axis. It means that you are trying to use the method `swaplevel()` on a flat index. You need to check what is the index by: ```python df.index ``` ## Step 6: IndexError: Too many levels: Index has only 3 levels, not 4 If we try to pass higher level than the one existing on the MultiIndex we will get an error: > IndexError: Too many levels: Index has only 3 levels, not 4 We need to check the levels with: ```python df.index.names len(df.index.names) ``` This will give us the names of the levels( and the number): ``` FrozenList([None, None, None]) 3 ``` or ```python df.index.shape ``` This will give us: ``` (4,) ``` Which is the number of the rows. ## Conclusion This article covers how to swap levels of MultiIndex. It shows different options and most common problems. It gives hints on how to deal with MultiIndex and how to get information about the levels. ### How to Read XML File with Python and Pandas URL: https://datascientyst.com/read-xml-file-python-pandas/ Last updated: 2022-10-13T20:24:36.000Z In this quick tutorial, we'll cover how to **read or convert XML file to Pandas DataFrame or Python data structure**. Since version 1.3 **Pandas offers an elegant solution for reading XML files**: `pd.read_xml()`. The short solutions is: ```python df = pd.read_xml('sitemap.xml') ``` With the single line above we **can read XML file to Pandas DataFrame or Python structure**. Below we will cover multiple examples in greater detail by using two ways: - `pd.read_xml()` - `xmltodict` \- Python library ## Setup Suppose we have simple XML file with the following structure: ```xml https://example.com/item-1 2022-06-02T00:00:00Z weekly https://example.com/item-2 2022-06-02T11:34:37Z weekly https://example.com/item-3 2022-06-03T19:24:47Z weekly ``` which we would like to read as Pandas DataFrame like shown below: | | loc | lastmod | changefreq | | - | -------------------------- | -------------------- | ---------- | | 0 | https://example.com/item-1 | 2022-06-02T00:00:00Z | weekly | | 1 | https://example.com/item-2 | 2022-06-02T11:34:37Z | weekly | | 2 | https://example.com/item-3 | 2022-06-03T19:24:47Z | weekly | or getting the links as Python list: ``` ['https://example.com/item-1', 'https://example.com/item-2', 'https://example.com/item-3'] ``` ## Step 1: Read XML File with read\_xml() The official documentation of method `read_xml()` is placed on this link: [pandas.read\_xml](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fxml.html?ref=datascientyst.com). To **read the local XML file in Python** we can give the absolute path of the file: ```python import pandas as pd df = pd.read_xml('sitemap.xml') ``` The result will be: | | loc | lastmod | changefreq | | - | -------------------------- | -------------------- | ---------- | | 0 | https://example.com/item-1 | 2022-06-02T00:00:00Z | weekly | | 1 | https://example.com/item-2 | 2022-06-02T11:34:37Z | weekly | | 2 | https://example.com/item-3 | 2022-06-03T19:24:47Z | weekly | The method has several useful parameters: - `xpath` \- The XPath to parse the required set of nodes for migration to DataFrame. - `elems_only` \- Parse only the child elements at the specified xpath. By default, all child elements and non-empty text nodes are returned. - `names` \- Column names for DataFrame of parsed XML data. - `encoding` \- Encoding of XML document. - `namespaces` \- The namespaces defined in XML document as dicts with key being namespace prefix and value the URI. ![](https://datascientyst.com/content/images/2022/06/read-xml-file-python-pandas.png) ## Step 2: Read XML File with read\_xml() - remote Now let's use **Pandas to read XML** from a remote location. The first parameter of `read_xml()` is: `path_or_buffer` described as: > String, path object (implementing os.PathLike\[str\]), or file-like object implementing a read() function. The string can be any valid XML string or a path. The string can further be a URL. Valid URL schemes include http, ftp, s3, and file. So we can **read remote files** the same way: ```python import pandas as pd df = pd.read_xml( f'https://s3.example.com/sitemap.xml.gz') ``` The final output will be exactly the same as before - DataFrame which has all values from the XML data. ## Step 3: Read XML File as Python list or dict Now suppose you need to **convert XML file to Python list or dictionary**. We need to read the XML file first, then to convert the file to DataFrame and finally to get the values from this DataFrame by: **Example 1: List** ```python df['loc'].to_list() ``` result: ``` ['https://example.com/item-1', 'https://example.com/item-2', 'https://example.com/item-3'] ``` **Example 2: Dictionary** ```python df['loc'].to_dict() ``` result: ``` {0: 'https://example.com/item-1', 1: 'https://example.com/item-2', 2: 'https://example.com/item-3'} ``` **Example 3: Dictionary - orient index** ```python df[['loc', 'changefreq']].to_dict(orient='index') ``` result: ``` {0: {'loc': 'https://example.com/item-1', 'changefreq': 'weekly'}, 1: {'loc': 'https://example.com/item-2', 'changefreq': 'weekly'}, 2: {'loc': 'https://example.com/item-3', 'changefreq': 'weekly'}} ``` ## Step 4: Read multiple XML Files in Python Finally let's see how to **read multiple identical XML files with Python** and Pandas. Suppose that files are identical with the following format: > [https://s3.example.com/sitemap{i}.xml.gz](https://s3.example.com/sitemap%7Bi%7D.xml.gz?ref=datascientyst.com) So: - [https://s3.example.com/sitemap1.xml.gz](https://s3.example.com/sitemap1.xml.gz?ref=datascientyst.com) - [https://s3.example.com/sitemap2.xml.gz](https://s3.example.com/sitemap2.xml.gz?ref=datascientyst.com) We can use the following code to read all files in a given range and concatenate them into a single DataFrame: ```python import pandas as pd df_temp = [] for i in (range(1, 10)): s = f'https://s3.example.com/sitemap{i}.xml.gz' df_site = pd.read_xml(s) df_temp.append(df_site) ``` The result is a list of DataFrames which can be concatenated into a single one by: ```python df_all = pd.concat(df_temp) ``` Now we have information from all XML files into df\_all. ## Step 5: Read XML File - xmltodict There is an **alternative solution for reading XML file in Python** by using the library: `xmltodict`. It can be installed by: ```python pip install xmltodict ``` To read XML file we can do: ```python import xmltodict with open('sitemap.xml') as fd: doc = xmltodict.parse(fd.read()) ``` The result is a dict: ``` {'urlset': {'@xmlns': 'http://www.sitemaps.org/schemas/sitemap/0.9', 'url': [{'loc': 'https://example.com/item-1', 'lastmod': '2022-06-02T00:00:00Z', 'changefreq': 'weekly'}, {'loc': 'https://example.com/item-2', 'lastmod': '2022-06-02T11:34:37Z', 'changefreq': 'weekly'}, {'loc': 'https://example.com/item-3', 'lastmod': '2022-06-03T19:24:47Z', 'changefreq': 'weekly'}]}} ``` Accessing elements can be done by: ```python doc['urlset']['url'][0] ``` the result is: ``` {'loc': 'https://example.com/item-1', 'lastmod': '2022-06-02T00:00:00Z', 'changefreq': 'weekly'} ``` ## Conclusion In this article, we covered several ways to **read XML file with Python and Pandas**. Now we know how to **read local or remote XML files**, using two Python libraries. Different options and parameters make the **XML conversion with Python - easy and flexible.** ### 136-read-xml URL: https://datascientyst.com/136-read-xml/ Last updated: 2022-06-07T07:30:51.000Z read\_xml() ### Pandas Cheat Sheet for Data Science URL: https://datascientyst.com/pandas-cheat-sheet-for-data-science/ Last updated: 2023-03-01T05:21:33.000Z A handy **Pandas Cheat Sheet** useful for the aspiring data scientists and contains ready-to-use codes for data wrangling. The **cheat sheet summarize the most commonly used Pandas features and APIs**. This cheat sheet will act as a **crash course for Pandas beginners** and help you with various fundamentals of Data Science. It can be used by experienced users as a quick reference. - [Pandas API Reference](https://pandas.pydata.org/pandas-docs/stable/reference/index.html?ref=datascientyst.com#api) - [Pandas User Guide](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/index.html?ref=datascientyst.com#user-guide) - [Data Wrangling with Pandas Cheat Sheet](https://pandas.pydata.org/Pandas%5FCheat%5FSheet.pdf?ref=datascientyst.com) ## Pandas setup Pandas is open source package for data science/data analysis ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_basics.png) `pip install pandas` Install Pandas library `import pandas as pd` Import Pandas `df = (pd.melt(df) .rename(columns={'variable': 'var'}) .query('val >= 200') )` Method Chaining `df[(df['col_1']=='a')&(df['col_2']>=10)]` Pandas syntax ## Data Structures Pandas has two data structures: Series (1 dimension) and DataFrame (multidimensional) ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_data_structures-1.png) ` s = pd.Series(['a', 'b', 'c'], index=[0 , 1, 2])` Create new Series `df = pd.DataFrame( {'col_1': [11, 12, 13], 'col_2': [21, 22, 23], 'col_3': [31, 32, 33]}, index=[0, 1, 2])` Create new DataFrame ## Read Import data from CSV, Excel, JSON, SQL, HTML, web ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_read.png) `pd.read_csv(filename)` From a CSV file `pd.read_csv(filename, header=None, nrows=5)` From a CSV file with parameters `pd.read_excel(filename)` From an Excel file `pd.read_sql(query, connection_object)` Reads from a SQL table/database `pd.read_json(json_string)` Reads from a JSON formatted string, URL or file. ## Write Write data to CSV, Excel, JSON, HTML ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_write.png) `df.to_csv(filename)` Writes to a CSV file `df.to_excel(filename)` Writes to an Excel file `df.to_json(filename)` Writes to a file in JSON format `df.to_html(filename)` Saves as an HTML table ## Inspect Data View stats, samples and summary of the data ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_info.png) `df.head(n)` First n rows `df.tail(n)` Last n rows `df.shape` Number of rows and columns `df.info()` Index, Datatype and Memory information `df.describe()` Summary statistics for numerical columns `s.value_counts(dropna=False)` (Series) Views unique values and counts `df.sample(n)` Randomly select n rows. `df.nlargest(n, 'col_1')` Select and order top n entries for column `df.nsmallest(n, 'col_1')` Select and order bottom n entries. `df.quantile([0.25,0.75])` Quantiles of each object ## Select Select data by index, by label, get subset ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_select.png) `s.loc[0]` (Series) Select by index `s.iloc[0]` (Series) Select by position `df['col_1']` Get single column as Series `df[['col_1', 'col_2']]` Get multiple columns as a DataFrame `df.iloc[0,:]` Select first row from DataFrame `df.iloc[0,0]` First element of first column `df.loc[df['col_1'] > 10, ['col_1', 'col_2']]` Select rows meeting logical condition, and only the specific columns `df.iat[1, 2] ` Access single value by index `df.at[3, 'col_2']` Access single value by label ## Add rows/columns Add new values to existing DataFrame ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_add.png) `df['new col'] = df['col'] * 100` Add new column based on other column `df['new col'] = False` Add new column single value `df.loc[-1] = [1, 2, 3]` Add new row at the end of DataFrame `df.append(df2, ignore_index = True)` add rows from DataFrame to existing DataFrame ## Drop rows/columns/nan Drop data from DataFrame ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_drop.png) `s.drop([0, 1])` (Series) Drop values from Series by index (row axis) `df.drop('col_1' , axis=1) ` Drop column by name col\_1 (column axis) `df.dropna()` Drops all rows that contain null values `df.dropna(axis=1)` Drops all columns that contain null values `df.dropna(axis=1,thresh=n)` Drops all rows have have less than n non null values ## Sort values/index Sort and rank values/index by one or multiple criteria ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_sort.png) `df.sort_values(by='col_1', ascending=False)` Sort values by column, ascending order `df.sort_values(by=['col_1', 'col_2'])` Sort values by columns `df.sort_index(ascending=False)` Sort object by labels (along an axis) in descending order `df.sort_values(by=[('col_1', 'col_2')])` Sort multindex by multiple levels `df.reset_index()` Reset the index of the DataFrame, moving index to columns ## Filter Filter data based on multiple criteria ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_filter.png) `df[df['col_1'] > 100]` Values greater than X `df[(df['col_1']=='a')&(df['col_2']>=10)]` Filter Multiple Conditions - & - and; | - or `df[df['date'] > '2022-02-22']` Date filtering `df[df['date'].dt.month == 2]` Filter with dt attributes `df[df['col_1'].str.contains('pan*', regex=True)]` Filter by regex `df[df['col_1'].isin(['pan', 'das'])]` Filter based on list of values `df.query('col_1 > 100')` Filter by queries `df.query('col_1 > 100 and col_2 = 0')` Filter by multiple queries ## Group by Group by and summarize data ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_groupby.png) `df.groupby('col_1')` Group by single column - return pandas.core.groupby.DataFrameGroupBy `df.groupby(['col_1', 'col_2'])` Group by multiple columns `df.groupby('col_1').groups` View groups `df.groupby('col_1').get_group(1)` Get group `df.groupby('col_1').count()` Get count per groups `df.groupby('col_1').agg([np.sum, np.mean])` Apply multiple agg functions on group `df.groupby('col_1').filter(lambda x: len(x) >= 5)` Filter groups `df.groupby('col_1').agg('count')` Aggregate group using function. `df.groupby('col_1').rank(method='dense')` Compute numerical data ranks (1 through n) along axis. ## Convert Convert to date, string, numeric ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_convert.png) `df['points'].astype(str)` convert to string `df['col_1'].astype('int64')` convert to int64 `df['col_1'].astype(float)` convert to float `pd.to_numeric(df['col_1'], errors='coerce')` convert to numeric `pd.to_datetime(df['date'], format='%Y-%m-%d')` convert string to date `pd.DataFrame(df['Values'].tolist(), columns=['col_1', 'col_1'])` Split column list to multiple columns `df['col_1'].apply(pd.Series)` Expand Series of dictionnaries ## Merge & Concat Merging, joniing and concatenating 2 and more DataFrames ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_merge.png) `df1.append(df2)` Adds the rows in df1 to the end of df2 (columns should be identical) `pd.concat([df1, df2],axis=1)` Adds the columns in df1 to the end of df2 (rows should be identical) `df1.join(df2,on=col1,how='inner')` SQL-style joins the columns in df1 with the columns on df2 where the rows for col have identical values. how can be one of 'left', 'right', 'outer', 'inner' ## Apply Applying functions to a column or DataFrame; lambda functions ![](https://datascientyst.com/content/images/2022/04/pandas_cheat_sheet_apply.png) `def calc(x): return x + 1 df.apply(calc, axis = 1)` Apply function to DataFrame `df[['col_1','col_2']].apply(calc)` Apply to multiple columns `df.apply(lambda x: x * -1 if x < 0 else x)` Apply lambda ![](https://datascientyst.com/content/images/2022/06/pandas_cheatsheet.png) ### Data Cleaning Steps with Python and Pandas URL: https://datascientyst.com/data-cleaning-steps-python-example/ Last updated: 2023-04-16T05:00:17.000Z Often we may need to **clean the data using Python and Pandas**. This tutorial explains the **basic steps for data cleaning by example**: - Basic exploratory data analysis - Detect and remove missing data - Drop unnecessary columns and rows - Detect outliers - Inconsistent data - Irrelevant features ## What is Data Cleaning? What is dirty Data? First let's see what is dirty data: > dirty data is inaccurate, incomplete or inconsistent data The common features of dirty data are: - spelling or punctuation errors - incorrect data associated with a field - incomplete data - outdated data - duplicated records The process of fixing all issues above is known as **data cleaning or data cleansing**. Usually **data cleaning process has several steps**: - normalization (optional) - detect bad records - correct problematic values - remove irrelevant or inaccurate data - generate report (optional) At the end of the process data should be: - complete - up to date - accurate - correct - consistent - relevant - normalized Difference of **Tidy data vs clean data? Data Tidying vs Data Cleaning?** Data cleaning is related to data quality. Data tidying is related to data structure. For tidy data - each observation is saved in its own row - each variable is saved in its own column ## Setup In this post we will use data from Kaggle - [A Short History of the Data-science](https://www.kaggle.com/code/softhints/a-short-history-of-the-data-science/data?ref=datascientyst.com). Above you can find a notebook related to [2019 Kaggle Machine Learning & Data Science Survey](https://www.kaggle.com/c/kaggle-survey-2019?ref=datascientyst.com). To read the data you need to use the following code: ```python import kaggle link = 'eswarankrishnasamy/2019-kaggle-machine-learning-data-science-survey' kaggle.api.authenticate() kaggle.api.dataset_download_file(link, file_name='multiple_choice_responses.csv', path='data/') ``` The downloaded data can be ready by: ```python import pandas as pd pd.read_csv('data/multiple_choice_responses.csv.zip', low_memory=False) ``` Data looks like: | Time from Start to Finish (seconds) | Q1 | Q2 | Q2\_OTHER\_TEXT | Q3 | | ----------------------------------- | --------------------------- | -------------------------------------- | ----------------------------------------------------- | ----------------------------------------- | | Duration (in seconds) | What is your age (# years)? | What is your gender? - Selected Choice | What is your gender? - Prefer to self-describe - Text | In which country do you currently reside? | | 510 | 22-24 | Male | \-1 | France | | 423 | 40-44 | Male | \-1 | India | | 83 | 55-59 | Female | \-1 | Germany | | 391 | 40-44 | Male | \-1 | Australia | ## Step 1: Exploratory data analysis in Python and Pandas To start we can do **basic exploratory data analysis in Pandas.** This will show us more about data: - data types - shape and size - missing values - sample data The first method is `head()` \- which returns the first 5 rows of the dataset. To see the first 5 rows and 5 columns we can do: `df.iloc[0:5,0:5]` ```python df.head() ``` The result is truncated for the first 5 columns: | | Time from Start to Finish (seconds) | Q1 | Q2 | Q2\_OTHER\_TEXT | Q3 | | - | ----------------------------------- | --------------------------- | -------------------------------------- | ----------------------------------------------------- | ----------------------------------------- | | 0 | Duration (in seconds) | What is your age (# years)? | What is your gender? - Selected Choice | What is your gender? - Prefer to self-describe - Text | In which country do you currently reside? | | 1 | 510 | 22-24 | Male | \-1 | France | | 2 | 423 | 40-44 | Male | \-1 | India | | 3 | 83 | 55-59 | Female | \-1 | Germany | | 4 | 391 | 40-44 | Male | \-1 | Australia | Next we can see information about the number of the columns and rows by `df.shape`: ```python df.shape ``` The result is a tuple showing 19718 rows and 246 columns: ``` (19718, 246) ``` Similar information we can get by `df.info()`: ```python df.info() ``` result: ``` RangeIndex: 19718 entries, 0 to 19717 Columns: 246 entries, Time from Start to Finish (seconds) to Q34_OTHER_TEXT dtypes: object(246) memory usage: 37.0+ MB ``` Finally we can get more details information about the data values by method `describe()`. This method will generate descriptive statistics (summarize the central tendency, dispersion and shape of a dataset’s distribution, excluding NaN values). ```python df.describe() ``` | | Time from Start to Finish (seconds) | Q1 | Q2 | Q2\_OTHER\_TEXT | Q3 | | ------ | ----------------------------------- | ----- | ----- | --------------- | ----- | | count | 19718 | 19718 | 19718 | 19718 | 19718 | | unique | 4169 | 12 | 5 | 46 | 60 | | top | 450 | 25-29 | Male | \-1 | India | | freq | 42 | 4458 | 16138 | 19668 | 4786 | ## Step 2: First rows as header read\_csv in Pandas So far we saw that the first row contains data which belongs to the header. We need to change how we read the data with `header=[0,1]`: ```python df = pd.read_csv('data/multiple_choice_responses.csv.zip', low_memory=False, header=[0,1]) ``` The above will [read the multiline header from the CSV file](https://datascientyst.com/read-excel-csv-multiple-line-headers-using-pandas/). In order to simplify the reading of the data we can [drop single level from the multi-index by](https://datascientyst.com/pandas-drop-multiindex-level/): ```python df.droplevel(level=1, axis=1) ``` ## Step 3: Data tidying in Pandas Next we can do **data tidying** because tidy data helps Pandas's vectorized operations. For example column 'Q1' looks like - we need to use the multi-index in order to read the column: ```python df[('Q1', 'What is your age (# years)?')] ``` resulted data is: ``` 0 22-24 1 40-44 2 55-59 3 40-44 4 22-24 ``` Can we split that into two columns? It looks like that all values are two numbers separated by '-' hyphen. The best is to confirm that observation by: ```python df[('Q1', 'What is your age (# years)?')].value_counts() ``` The last rows shows us one record - 70+ which needs special attention ``` 45-49 949 50-54 692 55-59 422 60-69 338 70+ 100 dtype: int64 ``` If we perform split operation on rows containing 70+ will result into: ```python df['Q1'].str.split('-', expand=True) ``` output ``` 0 70+ 1 None Name: 182, dtype: object ``` ![](https://datascientyst.com/content/images/2022/03/comicspin_data_cleaning.png) ## Step 4: Correcting and replacing data in Pandas Next we can see how to correct the data above. We can do **data correction** of cases 70+ in two ways: - replace value of 70+ with something else - split the column and fill the NaN values ### 4.1\. Replace values in column - Pandas To replace the values in the column we can use method `.str.replace('70+', '70-120', regex=False)` as follows: ```python df['Q1'].str.replace('70+', '70-120', regex=False) ``` ### 4.2\. Fill NaN with string or 0 - Pandas The other option is to fill the missing values after the split by: ```python df['max_age'].fillna(120) ``` we suppose that after the split we created new column 'max\_age' ## Step 5: Detect NaN values in column Pandas Now let's see how we can detect NaN values. This will help us **drop columns with NaN values**. ### 5.1 Columns which contains only NaN values To find columns which has only NaN values we can use two methods: - `isna()` - `all()` ```python s = df.isna().all() ``` This will give a new Series with column name and True or False - depending on the NaN values. If a column has only NaN values we will get True. To find columns which contain NaN values we can use: ```python s[s == True] ``` the result is: ``` Series([], dtype: bool) ``` There's no column which contains only NaN values ### 5.2 Detect columns with NaN values To **detect columns which has NaN values** we can use: ```python df.isna().any() ``` This will result into: ``` Time from Start to Finish (seconds) False Q1 False Q2 False Q2_OTHER_TEXT False Q3 False ... Q34_Part_9 True Q34_Part_10 True Q34_Part_11 True Q34_Part_12 True Q34_OTHER_TEXT False Length: 246, dtype: bool ``` So columns like 'Q34\_Part\_9' have NaN values. Columns like 'Q1' don't have NaN values. ## Step 6: Drop columns in Pandas Let's say that we would like to **drop columns based on name or NaN values**. We can do that in several ways: ### 6.1 Drop one column by name Parameters needed to drop columns are `axis=1` and `inplace=True` \- which means that operation will affect DataFrame. ```python df.drop('Q1', axis=1, inplace=True) ``` ### 6.2 Drop multiple columns by name We can list several column which to be removed by: ```python df.drop(['Q1', 'Q2'], axis=1, inplace=True) ``` ### 6.3 Drop columns with NaN values Finally we can **drop columns which has NaN values**: ```python df.dropna(axis=1, how='any') ``` We can use parameters like: - `how` \- 'all' or 'any' - `subset` \- list of columns - `tresh` \- the number of NaN values required to remove the column - `inplace` ## Step 7: Detect and drop duplicate rows in Pandas To **detect duplicate values in the DataFrame** we can use the method `duplicated()`. To detect duplicate rows in Pandas DataFrame we can use: ```python df[df.duplicated()] ``` This results in 4 duplicated rows: ``` 4 rows × 246 columns ``` We can use parameter: `keep` - first - only the first rows - last - the last row only - False - show all rows For example get indexes of all detected duplications: ```python df[df.duplicated(keep=False)].index ``` > Int64Index(\[11228, 12344, 16413, 16547, 16653, 18705, 19258, 19705\], dtype='int64') Since we have 246 columns (answers) it's pretty suspicious that there are full duplications. We can use method `df.drop_duplicates(subset=['Q1'])` in order to drop duplicated rows in Pandas: ```python df.drop_duplicates(subset=['Q1', 'Q2']) ``` ## Step 8: Detect outliers in Pandas We can **detect outliers in Pandas** in many ways. Here we will cover basic detection of numeric data: - check stats for the column - min, max and percentiles - visually by plotting values Suppose we work with column: 'Time from Start to Finish (seconds)' We can see the min, max and the percentiles by: ```python df['Time from Start to Finish (seconds)'].describe() ``` The result is: ``` count 19717.000000 mean 14341.281027 std 74166.106601 min 23.000000 25% 340.000000 50% 540.000000 75% 930.000000 max 843612.000000 Name: (Time from Start to Finish (seconds), Duration (in seconds)), dtype: float64 ``` So we have time for the survey from 23 up to 843612 seconds. Probably we can exclude some of them. Another way to detect outliers is visually by plotting data like: ![data-cleaning-steps-python-detect_outliers](https://datascientyst.com/content/images/2022/03/data-cleaning-steps-python-detect_outliers.png) From the image above we can decide what is the threshold which makes sense for us. ## Step 9: Detect errors, typos and misspelling in Pandas Finally let's check how we can **detect typos and misspelled words in Pandas DataFrame**. This will show how we can work with inconsistent or incomplete data. For this purpose we are going to read file - 'other\_text\_responses.csv' which will be `df_other`. The reason is that it contains free text input. Let's read the third column of this DataFrame by: ```python df_other[df_other.columns[3]].value_counts().head(10) ``` The result is: ``` Excel 865 Microsoft Excel 392 excel 263 MS Excel 67 Google Sheets 61 Google sheets 44 Microsoft excel 38 Excel 33 microsoft excel 27 EXCEL 25 ``` We can see different variations of the same tool - Excel. In order to detect similar values we will use Python library `difflib`: ```python import difflib difflib.get_close_matches('excl', ['Excel', 'Microsoft Excel ', 'MS Excel', 'excel'], n=1, cutoff=0.7) ``` The result of this will be: ``` ['excel'] ``` So we can use Python in order to detect and fix misspelled words. Code like the one below can help us create new column with corrected values: ```python import difflib correct_values = {} words = df_other["Q14_Part_3_TEXT"].value_counts(ascending=True).index for keyword in words: similar = difflib.get_close_matches(keyword, words, n=20, cutoff=0.6) for x in similar: correct_values[x] = keyword df_other["corr"] = df_other["Q14_Part_3_TEXT"].map(correct_values) ``` ## Conclusion In this article, we learned **what is clean data and how to do data cleaning in Pandas and Python**. Some topics which we discussed are NaN values, duplicates, drop columns and rows, outlier detection. We saw all the steps of the data cleaning process with examples. We covered important topics like tidy data and data quality. ### How to Solve Error Tokenizing Data on read_csv in Pandas URL: https://datascientyst.com/solve-error-tokenizing-data-read_csv-pandas/ Last updated: 2023-11-18T23:53:12.000Z In this tutorial, we'll see **how to solve a common Pandas read\_csv() error – Error Tokenizing Data.** The full error is something like: > ParserError: Error tokenizing data. C error: Expected 2 fields in line 4, saw 4 The **Pandas parser error when reading csv** is very common but difficult to investigate and solve for big CSV files. There could be many different reasons for this error: - "wrong" data in the file - different number of columns - mixed data - several data files stored as a single file - nested separators - wrong parsing parameters for read\_csv() - different separator - line terminators - wrong parsing engine Let's cover the steps to diagnose and solve this error ## 1: Reason for pandas.parser.CParserError: Error tokenizing Suppose we have CSV file like: ``` col_1,col_2,col_3 11,12,13 21,22,23 31,32,33,44 ``` which we are going to read by - `read_csv()` method: ```python import pandas as pd pd.read_csv('test.csv') ``` We will get an error: ``` ParserError: Error tokenizing data. C error: Expected 3 fields in line 4, saw 4 ``` We can easily see where the problem is. But what should be the solution in this case? Remove the 44 or add a new column? It depends on the context of this data. If we don't need the bad data we can use parameter - `on_bad_lines='skip'` in order to skip bad lines: ```python pd.read_csv('test.csv', on_bad_lines='skip') ``` For older Pandas versions you may need to use: `error_bad_lines=False` which will be deprecated in future. If we like to get warning messages for the problematic lines then **on\_bad\_lines='warn'** can be used. Using `warn` instead of `skip` will produce: ```python pd.read_csv('test.csv', on_bad_lines='warn') ``` warning like: ``` b'Skipping line 4: expected 3 fields, saw 4\n' ``` To find more about how to [drop bad lines with read\_csv()](https://datascientyst.com/drop-bad-lines-with-read%5Fcsv-pandas/) read the linked article. ![](https://datascientyst.com/content/images/2022/03/solve-error-tokenizing-data-read_csv-pandas.png) ## 2: How To Fix pandas.parser.CParserError: Error tokenizing In some cases the reason could be the separator used to read the CSV file. In this case we can open the file and check its content. If you don't know how to [read huge files in Windows or Linux](https://softhints.com/how-to-open-and-search-huge-files-windows-linux/?ref=datascientyst.com) \- then check the article. Depending on your OS and CSV file you may need to use different parameters like: - `sep` \- - `lineterminator` \- - `engine` More information on the parameters can be found in [Pandas doc for read\_csv()](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) If the **CSV file has tab as a separator and different line endings** we can use: ```python import pandas as pd pd.read_csv('test.csv', sep='\t', lineterminator='\r\n') ``` Note that `delimiter` is an alias for `sep`. ## 3: Use different engine for read\_csv() The default C engine cannot automatically detect the separator, but the `python` parsing engine can. There are 3 engines in the latest version of Pandas: - `c` - `python` - `pyarrow` Python engine is the slowest one but the most feature-complete. **Using `python` as engine can help detecting the correct delimiter** and solve the Error Tokenizing Data: ```python pd.read_csv('test.csv', engine='python') ``` ## 4: Identify the headers to solve Error Tokenizing Data **Sometimes Error Tokenizing Data problem may be related to the headers**. For example multiline headers or additional data can produce the error. In this case we can skip the first or last few lines by: ```python pd.read_csv('test.csv', skiprows=1, engine='python') ``` or skipping the last lines: ```python pd.read_csv('test.csv', skipfooter=1, engine='python') ``` Note that we need to use `python` as the `read_csv()` engine for parameter - `skipfooter`. In some cases `header=None` can help in order to solve the error: ```python pd.read_csv('test.csv', header=None) ``` ## 5: Autodetect skiprows for read\_csv() Suppose we had CSV file like: ``` Dates 2022-03-20 2022-03-21 2022-03-22 Raw data col_1,col_2,col_3 11,12,13 21,22,23 31,32,33 ``` We are interested in reading the data stored after the line "Raw data". If this line is changing for different files and we need to autodetect the line. To search for a line containing some match in Pandas we can use a separator which is not present in the file "@@". In case of multiple occurrences we can get the biggest index. **To autodetect skiprows parameter in Pandas `read_csv()` use**: ```python df_temp = pd.read_csv('test.csv', sep='@@', engine='python', names=['col']) ix_last_start = df_temp[df_temp['col'].str.contains('# Raw data')].index.max() ``` result is: ``` 4 ``` Finding the starting line can be done by visually inspecting the file in the text editor. Then we can simply read the file by: ```python df_analytics = pd.read_csv(file, skiprows=ix_last_start) ``` ## 6: Update CSV file to align headers and values As we mentioned the most common reason for this error is the different number of values in comparison to columns. The example CSV file will raise the error: ``` col_1,col_2,col_3 1,2,3 4,5,6 7,8,9,0 ``` Easy fix for **ParserError: Error tokenizing data. C error: Expected 3 fields in line 4, saw 4** is to change the headers of the CSV file to the number shown in the error. So you can simply change the file to: ``` col_1,col_2,col_3,col_4 1,2,3 4,5,6 7,8,9,0 ``` This will fix the Pandas error - **ParserError: Error tokenizing data** ## Conclusion In this post we saw how to investigate the error: ``` ParserError: Error tokenizing data. C error: Expected 3 fields in line 4, saw 4 ``` We covered different reasons and solutions in order to read any CSV file with Pandas. Some solutions will warn and skip problematic lines. Others will try to resolve the problem automatically. ### How to Remove Trailing and Consecutive Whitespace in Pandas URL: https://datascientyst.com/remove-trailing-consecutive-whitespace-pandas/ Last updated: 2022-03-21T14:58:34.000Z In this short guide, we'll see how to **remove consecutive, leading and trailing whitespaces in Pandas**. We will cover how to strip all spaces in: - entire DataFrame - multiple columns - columns names - read\_csv and whitespaces Below we can find the steps to follow for cleaning whitespaces in Pandas. ## Setup In the post, we'll use the following DataFrame, which consists of several rows and columns: ```python import pandas as pd data = {'title': [" The jungle book", "Beautiful mind ", " Rose "], 'number': [ " 123 456 789", " 234 567 890 ", " 1 2 3 " ]} df = pd.DataFrame.from_dict(data) ``` DataFrame looks like: | | title | number | | - | --------------- | ----------- | | 0 | The jungle book | 123 456 789 | | 1 | Beautiful mind | 234 567 890 | | 2 | The Rose | 1 2 3 | Sometimes white spaces are not visible in the DataFrame because of the way browsers work. Consecutive spaces are not shown by most browsers. The real printed data will look like: ``` title number 0 The jungle book 123 456 789 1 Beautiful mind 234 567 890 2 The Rose 1 2 3 ``` ## Step 1: Strip leading and trailing whitespaces in Pandas To **strip whitespaces from a single column in Pandas** We will use method - [str.strip()](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.strip.html?ref=datascientyst.com). This method is applied on a single column a follows: ```python df['title'].str.strip() ``` After the operation the **leading and trailing whitespace are trimmed**: ``` 0 The jungle book 1 Beautiful mind 2 The Rose Name: title, dtype: object ``` **Note:** strip() only works on leading and trailing white space. ## Step 2: Replace consecutive spaces in Pandas To **remove consecutive spaces from a column in Pandas** we will use method `replace` with `regex=True`: ```python df['title'].replace(r'\s+',' ', regex=True) ``` The result is: ``` 0 The jungle book 1 Beautiful mind 2 The Rose Name: title, dtype: object ``` We can notice that leading and trailing whitespace are still present. To remove them we need to chain both methods: ```python df['title'].str.strip().replace(r'\s+',' ', regex=True) ``` The final result is: ``` 0 The jungle book 1 Beautiful mind 2 The Rose Name: title, dtype: object ``` Consecutive, leading and trailing whitespace are removed from the column. ![](https://datascientyst.com/content/images/2022/03/remove-trailing-consecutive-whitespace-pandas.png) To **remove whitespaces from the column names** we can use the following syntax: ```python df.columns.str.strip() ``` ## Step 3: Remove whitespaces from multiple columns To **remove leading and trailing whitespace from multiple columns in Pandas** we can use the following code: ```python for column in ['title', 'number']: df[column] = df[column].apply(lambda x:x.strip()) ``` or to remove also the consecutive spaces we can use: ```python for column in ['title', 'number']: df[column] = df[column].apply(lambda x:x.strip().replace(r'\s+',' ') ) ``` ## Step 4: Remove whitespaces from entire DataFrame To **remove whitespaces from the entire DataFrame in Pandas** we will loop over all columns in the DataFrame by: ```python for col in df.select_dtypes('object'): df[col] = df[col].str.strip() ``` There are several ways to handle the possible errors. One way is to loop over the string columns - *dtype 'object'* or by catching the errors in try except block: ```python for col in df.columns: try: df[col] = df[col].str.strip() except AttributeError: pass ``` Depending on the need we can remove consecutive, leading or trailing whitespaces. ## Step 5: Skip whitespace with read\_csv() Finally we can **remove the whitespaces read by method `read_csv()`**. This can be done by parameters: - `delim_whitespace` \- Specifies whether or not whitespace (e.g. ' ' or ' ') will be used as the sep. Equivalent to setting sep='\\s+'. - `sep` \- regex can be used like - ',\\s+ - `skipinitialspace` \- Skip spaces after delimiter. ```python df = pd.read_csv(file, sep=',\s+', skipinitialspace=True) ``` This will help us skipping the spaces after delimiters. Combination of the above parameters might be needed in order to get optimal results. More can be found on the in the [read\_csv() documentation](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com). This will help removing the white space from the column names. ## Conclusion To summarize, in this article, we've seen examples of text cleaning. How to remove leading and trailing whitespace. How to use regex to remove consecutive spaces from data in one or multiple columns. And finally, we've seen how to manage errors and clean the whole DataFrame. ### How to Convert HTML to Text with Python and Pandas URL: https://datascientyst.com/convert-html-to-text-python-pandas/ Last updated: 2022-03-15T10:52:28.000Z In this short guide, we'll see how to **convert HTML to raw text with Python and Pandas**. It is also known as **text extraction from HTML tags**. ## 2\. Setup In this Python guide, we'll use the following DataFrame, which consists of two columns. Column `html` contains HTML tags and text inside the tags: ```python import pandas as pd data = {'title': ["Intro", "Code", "Code Text"], 'html': [ "Hello", "

Use applymap

", "

Code:

df.head()" ]} df = pd.DataFrame.from_dict(data) ``` DataFrame looks like: | | title | html | | - | --------- | ----------------------------------- | | 0 | Intro | Hello | | 1 | Code |

Use applymap

| | 2 | Code Text |

Code:

df.head() | We would like to **extract the raw text from the column without the HTML tags with Python**: | | html | | - | --------------- | | 0 | Hello | | 1 | Use applymap | | 2 | Code: df.head() | ## Step 1: Install Beautiful Soup library First we will need to install Python library - [beautifulsoup4](https://pypi.org/project/beautifulsoup4/?ref=datascientyst.com) by: ```python pip install beautifulsoup4 ``` The official documentation of the library is available on this link: [Beautiful Soup Documentation](https://www.crummy.com/software/BeautifulSoup/bs4/doc/?ref=datascientyst.com). The Beautiful Soup is described as: > Beautiful Soup is a Python library for pulling data out of HTML and XML files. ## Step 2: Extract text from HTML tags by Python Now let's check how we can **extract the text from HTML code or tags in Python.** We will demonstrate the extraction in simple example. Suppose we have the following HTML document: ```python html_doc = """The Dormouse's story

The Dormouse's story

Once upon a time there were three little sisters; and their names were

""" ``` We can extract the raw text from the HTML tags by: ```python from bs4 import BeautifulSoup soup = BeautifulSoup(html_doc, 'html.parser') print(soup.get_text()) ``` Which will result in: ``` The Dormouse's story The Dormouse's story Once upon a time there were three little sisters; and their names were ``` Or we can pretty print the HTML code by: ```python print(soup.prettify()) ``` the output is: ```html The Dormouse's story

The Dormouse's story

Once upon a time there were three little sisters; and their names were

``` ## Step 3: HTML to raw text in Pandas In order to **convert HTML to raw text we will apply BeautifulSoup library to Pandas column.** To apply the BeautifulSoup function `soup.get_text()` to Pandas column we can use the following code: ```python df[['html']].applymap(lambda text: BeautifulSoup(text, 'html.parser').get_text()) ``` The result is: | | html | | - | --------------- | | 0 | Hello | | 1 | Use applymap | | 2 | Code: df.head() | How does it work? We are applying the function `.get_text()` with `html.parser` to each row from the DataFrame - `df[['html']]` \- in this case it has only a single column. If we pass non HTML column or NaNs we will get errors. ## Conclusion In this post, we saw how to **convert HTML tags or document into raw text with Python and Pandas.** We've learn also how to apply **BeautifulSoup library function to Pandas DataFrame.** The image below shows all the steps which we covered: ![](https://datascientyst.com/content/images/2022/03/convert-html-to-text-python-pandas.png) ### How to Suppress and Format Scientific Notation in Pandas URL: https://datascientyst.com/suppress-format-scientific-notation-pandas/ Last updated: 2022-03-15T09:57:52.000Z **Pandas use scientific notation** to display float numbers. To **disable or format scientific notation in Pandas/Python** we will use option `pd.set_option` and other Pandas methods. All solutions were tested in Jupyter Notebook and JupyterLab. ## Setup In the post, we'll use the following DataFrame, which is created by the following code: ```python import pandas as pd data = {'int': [1, 2.0, 12], 'big int': [1, 2, 101548484845], 'big int + float': [1, 2.0, 101548484845], 'float': [0.1, 0.2, 0.003], 'big float': [0.1, 0.2, 0.000000025]} pd.DataFrame.from_dict(data) ``` DataFrame looks like: | int | big int | big int + float | float | big float | | ---- | ------------ | --------------- | ----- | ------------ | | 1.0 | 1 | 1.000000e+00 | 0.100 | 1.000000e-01 | | 2.0 | 2 | 2.000000e+00 | 0.200 | 2.000000e-01 | | 12.0 | 101548484845 | 1.015485e+11 | 0.003 | 2.500000e-08 | ## Step 1: Scientific notation in Pandas The DataFrame above consists of several columns named depending on the values. As we can see that some **float numbers cause Pandas to display numbers in scientific notation**. Big integer numbers with floats(in the same column) will be displayed in scientific notation too. The picture below shows both the scientific notation and the normal numbers. ![](https://datascientyst.com/content/images/2022/03/disable-format-scientific-notation-pandas.png) ## Step 2: Disable scientific notation in Pandas - globally We can **disable scientific notation in Pandas and Python by setting higher float precision**: ```python pd.set_option('display.float_format', lambda x: '%.9f' % x) ``` Now **float numbers will be displayed without scientific notation**: | int | big int | big int + float | float | big float | | ------------ | ------------ | ---------------------- | ----------- | ----------- | | 1.000000000 | 1 | 1.000000000 | 0.100000000 | 0.100000000 | | 2.000000000 | 2 | 2.000000000 | 0.200000000 | 0.200000000 | | 12.000000000 | 101548484845 | 101548484845.000000000 | 0.003000000 | 0.000000025 | to reset to the original settings we can use: ```python pd.reset_option('display.float_format', silent=True) ``` We need to provide the option which will be reset `'display.float_format'` ## Step 3: Format scientific notation in Pandas If we like to **change float number format in Pandas in order to suppress the scientific notation** than we can use - `lambda` in combination with formatter '%.9f': ```python df['big float'].apply(lambda x: '%.9f' % x) ``` Before the change we have: ``` 0 1.000000e-01 1 2.000000e-01 2 2.500000e-08 Name: big float, dtype: float64 ``` After the change we will have new float format: ``` 0 0.100000000 1 0.200000000 2 0.000000025 Name: big float, dtype: object ``` As we can see that type is changed from `float64` to `object`. This means that we could say that we (in some way) - **convert the scientific notation for float numbers to string**. ## Step 4: read\_csv and scientific notation in Pandas Sometimes we may need to use or not scientific notation for `read_csv` in Pandas. First **scientific notation for method `read_csv` depends mostly depends on parameter - `dtype`**. This will try to infer float numbers: ```python import pandas import numpy as np pandas.read_csv(filepath, dtype=np.float64) ``` Once data is read than we can convert or round it by: ```python df = df.apply(pd.to_numeric, errors='coerce') ``` or round float numbers with method `round`: ```python df.round(5) ``` This will change the scientific format only for float numbers: | int | big int | big int + float | float | big float | | ---- | ------------ | --------------- | ----- | --------- | | 1.0 | 1 | 1.000000e+00 | 0.1 | 0.1 | | 2.0 | 2 | 2.000000e+00 | 0.2 | 0.2 | | 12.0 | 101548484845 | 1.015485e+11 | 0.003 | 0.0 | ## Step 5: Format scientific notation to\_html If we like to disable or format scientific notation in Pandas method `to_html` then we should use parameter - `float_format`: ```python df.to_html(float_format='{:10.9f}'.format) ``` Now Pandas will generate Data with precision which will show the numbers without the scientific formatting. **Note:** `{:10.9f}` can be read as: - `10` \- specifies the total length of the number including the decimal portion - `9` \- is used to specify 9 decimal points Other examples: `{:30,.18f}` and `{:,.3f}` ## Conclusion In this post, we covered several examples when we work with **scientific notation in Pandas.** Now we can disable or change the format of float numbers shown in scientific notation. ### How to show all columns and rows in Pandas URL: https://datascientyst.com/pandas-show-all-columns-rows/ Last updated: 2022-10-28T06:14:54.000Z Hi all, in this tutorial, we'll learn how to **show all columns and rows in Pandas**. The solutions described in the article worked in Jupyter Notebooks and for printing in Python. ## Setup We'll use the following DataFrame from [Kaggle - IMDB 5000 Movie Dataset](https://www.kaggle.com/carolzhangdc/imdb-5000-movie-dataset?select=movie%5Fmetadata.csv&ref=datascientyst.com), which can be downloaded and read with Python/Pandas: ```python import kaggle import pandas as pd kaggle.api.authenticate() kaggle.api.dataset_download_file('carolzhangdc/imdb-5000-movie-dataset', file_name='movie_metadata.csv', path='data/') df = pd.read_csv('data/movie_metadata.csv.zip') ``` DataFrame output in JupyterLab limit the number of shown rows and columns: ![pandas-show-all-columns](https://datascientyst.com/content/images/2022/10/pandas-show-all-columns.webp) On the image we can see all rows and columns displayed by removing default thresholds. ## Step 1: Pandas show all columns - max\_columns By default Pandas will display only a limited number of columns. The limit depends on the usage. In this article you can learn more about the limits: [How to Show All Columns, Rows and Values in Pandas](https://softhints.com/pandas-display-all-columns-and-show-more-rows/?ref=datascientyst.com) To **show all columns in Pandas** we can set the option: `pd.option_context` \- `display.max_columns` to None. ```python with pd.option_context("display.max_columns", None): display(df) ``` This will show all columns in the current DataFrame. `display.max_columns` is described as: > In case Python/IPython is running in a terminal this is set to 0 by default and pandas will correctly auto-detect the width of the terminal and switch to a smaller format in case all columns would not fit vertically. ## Step 2: Pandas show all rows - max\_rows Pandas will show the first and last rows if the number of the rows is bigger than the threshold. All the other rows will be hidden. To **show all rows in Pandas** we can use option - `display.max_rows` equal to None or some other limit: ```python with pd.option_context("display.max_rows", None): display(df) ``` The option `max_rows` is described as: > This sets the maximum number of rows pandas should output when printing out various output. For example, this value determines whether the repr() for a DataFrame prints out fully or just a truncated or summary repr. ‘None’ value means unlimited. ×Be careful because showing or printing all rows/columns might cause performance issues in the browser. Even it can break the Jupyter Notebook. ## Step 3: Pandas set column width - max\_colwidth Values with length higher than 50 characters will be truncated when we print or display a big DataFrame. To **show full column width in Pandas** we can use: ```python with pd.option_context("display.max_colwidth", None): display(df) ``` ×Note that option max_colwidth works also for columns containing sequences. Definition of `max_colwidth` is: > The maximum width in characters of a column in the repr of a pandas data structure. When the column overflows, a “…” placeholder is embedded in the output. ‘None’ value means unlimited. ![pandas-show-all-columns-values](https://datascientyst.com/content/images/2022/03/pandas-show-all-columns-values.png) ## Step 4: Pandas show all columns and rows - pd.option\_context Finally if you like to **show all rows, columns and values from a Pandas DataFrame** we can use the following code to stop Pandas from hiding rows, columns and values: ```python import pandas as pd pd.set_option('display.max_rows', None) pd.set_option('display.max_columns', None) pd.set_option('display.width', None) pd.set_option('display.max_colwidth', None) ``` This will be applied for the whole Jupyter Notebook or Python code. So be careful otherwise we may face issues. ## Conclusion We saw how to **show all rows, columns and values in Pandas.** We've learned how to set options globally and for context. If you like to learn more about Pandas display option please refer: [Pandas Options and settings](https://pandas.pydata.org/docs/user%5Fguide/options.html?ref=datascientyst.com) ### How to Map Column with Dictionary in Pandas URL: https://datascientyst.com/pandas-map-column-dictionary/ Last updated: 2025-03-29T13:34:26.000Z ## 1\. Overview In this tutorial, we'll learn **how to map column with dictionary in Pandas DataFrame**. We are going to use Pandas method [pandas.Series.map](https://pandas.pydata.org/docs/reference/api/pandas.Series.map.html?ref=datascientyst.com) which is described as: > Map values of Series according to an input mapping or function. There are several different scenarios and considerations: - remap values in the same column - add new column with mapped values from another column - not found action - keep existing values Let's cover all examples in the next sections. The image below illustrates how to map column values work: ![](https://datascientyst.com/content/images/2022/01/how-to-map-new-column-from-dictionary-in-pandas-dataframe.png) ## 2\. Setup In the post, we'll use the following DataFrame, which consists of several rows and columns: ```python import pandas as pd import numpy as np data = {'Member': {0: 'John', 1: 'Bill', 2: 'Jim', 3: 'Steve'}, 'Disqualified': {0: 0, 1: 1, 2: 0, 3: 1}, 'Paid': {0: 1, 1: 0, 2: 0, 3: np.nan}} df = pd.DataFrame(data) ``` Data looks like: | | Member | Disqualified | Paid | | - | ------ | ------------ | ---- | | 0 | John | 0 | 1.0 | | 1 | Bill | 1 | 0.0 | | 2 | Jim | 0 | 0.0 | | 3 | Steve | 1 | NaN | ## 3\. Pandas map Column with Dictionary First let's start with the most simple case - **map values of column with dictionary**. We are going to use method - [pandas.Series.map](https://pandas.pydata.org/docs/reference/api/pandas.Series.map.html?ref=datascientyst.com). We are going to map column *Disqualified* to boolean values - 1 will be mapped as `True` and 0 will be mapped as `False`: ```python dict_map = {1: 'True', 0: 'False'} df['Disqualified'].map(dict_map) ``` The result is a new Pandas Series with the mapped values: ``` 0 False 1 True 2 False 3 True Name: Disqualified, dtype: object ``` ### 3.1 Map column values in DataFrame We can assign this result Series to the same column by: ```python df['Disqualified'] = df['Disqualified'].map(dict_map) ``` ### 3.2 Map dictionary to new column in Pandas To map dictionary from existing column to new column we need to change column name: ```python df['Disqualified Boolean'] = df['Disqualified'].map(dict_map) ``` **Note:** In case of a different DataFrame be sure that indices match ## 4\. Mapping column values and preserve values(NaN) What will happen if a value is not present in the mapping dictionary? In this case we will end with `NA` value: ```python df['Paid'].map(dict_map ) ``` result: ``` 0 True 1 False 2 NaN 3 NaN Name: Paid, dtype: object ``` In order to keep the not mapped values in the result Series we need to fill all missing values with the values from the column: ```python df['Paid'].map(dict_map).fillna(df['Paid']) ``` This will result into: ``` 0 True 1 False 2 3.0 3 NaN Name: Paid, dtype: object ``` To keep NaNs we can add parameter - `na_action='ignore'`: ```python df['Disqualified'].map(dict_map, na_action='ignore') ``` ## 5\. Map Column in Pandas - map() vs replace() An alternative solution to map column to dict is by using the function [pandas.Series.replace](https://pandas.pydata.org/docs/reference/api/pandas.Series.replace.html?ref=datascientyst.com). The syntax is similar but the result is a bit different: ```python df["Paid"].replace(dict_map) ``` In the result Series the original values of the column will be present: ``` 0 True 1 False 2 3.0 3 NaN Name: Paid, dtype: object ``` Another difference between functions map() and replace() are the parameters: - `.replace(dict_map, inplace=True)` \- applying changes on the Series itself - \`df\['Paid'\].map(dict\_map, na\_action='ignore') - to avoid applying the function to missing values (and keep them as NaN) Finally we can mention that `replace()` can be much slower in some cases. ## 6\. Map column with s.update() in Pandas Another option to **map values of a column based on a dictionary values** is by using method `s.update()` \- [pandas.Series.update](https://pandas.pydata.org/docs/reference/api/pandas.Series.update.html?ref=datascientyst.com) This can be done by: ```python df['Paid'].update(pd.Series(dict_map)) ``` The result will be update on the existing values in the column: ``` 0 False 1 True 2 3.0 3 NaN Name: Paid, dtype: object ``` The function is described as: > Modify Series in place using values from passed Series. > Uses non-NA values from passed Series to make updates. Aligns on index ## 7\. Map dictionary to new column in Pandas DataFrame Finally we can use pd.Series() of **Pandas to map dict to new column**. The difference is that we are going to use the index as keys for the dict: ```python df["Disqualified mapped"] = pd.Series(dict_map) ``` To use a given column as a mapping we can use it as an index. Then we an create the mapping by: ```python df = df.set_index(['Disqualified']) df['Disqualified mapped'] = pd.Series(dict_map) ``` ## 8\. Conclusion In this tutorial, we saw several options **to map, replace, update and add new columns based on a dictionary in Pandas**. We first looked into using the best option `map()` method, then how to keep not mapped values and NaNs, update(), replace() and finally by using the indexes. ## 9\. Resources - [Remap values in pandas column with a dict, preserve NaNs](https://stackoverflow.com/questions/20250771/remap-values-in-pandas-column-with-a-dict-preserve-nans?ref=datascientyst.com) ### How to Compare Two Pandas DataFrames and Get Differences URL: https://datascientyst.com/compare-two-pandas-dataframes-get-differences/ Last updated: 2022-10-28T06:11:07.000Z ## 1\. Overview In this tutorial, we're going to **compare two Pandas DataFrames side by side and highlight the differences**. We'll first look into Pandas method `compare()` to find differences between values of two DataFrames, then we will cover some advanced techniques to highlight the values and finally how to compare stats of the DataFrames. The image below show the final result: ![compare-two-pandas-dataframes-get-differences](https://datascientyst.com/content/images/2022/10/compare-two-pandas-dataframes-get-differences.webp) ## 2\. Setup Let's have the next DataFrame created by the code below: ```python import pandas as pd data = [('11','12','13'), ('21','22','23'), ('31','32','33')] df = pd.DataFrame(data, columns = ('col_1', 'col_2', 'col_3' )) ``` DataFrame looks like: | | col\_1 | col\_2 | col\_3 | | - | ------ | ------ | ------ | | 0 | 11 | 12 | 13 | | 1 | 21 | 22 | 23 | | 2 | 31 | 32 | 33 | Let's create second DataFrame by copying the first one and changing some values: ```python df2 = df.copy() df2.iloc[1,1] = 32 df2.iloc[2,1] = 22 df2.iloc[2,2] = 44 ``` The second DataFrame looks like: | | col\_1 | col\_2 | col\_3 | | - | ------ | ------ | ------ | | 0 | 11 | 12 | 13 | | 1 | 21 | 32 | 23 | | 2 | 31 | 22 | 44 | Can you find what are the differences just by watching both DataFrames? In next steps we will **compare two DataFrames in Pandas**. ## 3\. Compare Two Pandas DataFrames to Get Differences Pandas offers method: [pandas.DataFrame.compare](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.compare.html?ref=datascientyst.com) since version 1.1.0. It gives the **difference between two DataFrames** \- the method is executed on DataFrame and take another one as a parameter: ```python df.compare(df2) ``` The default result is new DataFrame which has differences between both DataFrames. The new DataFrame has multi-index - first level is the column name, the second one are the values from the both DataFrames which are compared: | | col\_2 | col\_3 | | | | - | ------ | ------ | ---- | ----- | | | self | other | self | other | | 1 | 22 | 32 | NaN | NaN | | 2 | 32 | 22 | 33 | 44 | The method has several parameters which we will cover in next sections: - `align_axis` - `keep_shape` - `keep_equal` ## 4\. Compare Two Pandas DataFrames Side by Side - keeping all values If you like to keep the original shape and all values (even if they are equal) then you can use - `keep_shape` and `keep_equal`. Keep the original shape - all equal values are replaced by NaN values: ```python df.compare(df2, keep_shape=True) ``` | | col\_1 | col\_2 | col\_3 | | | | | - | ------ | ------ | ------ | ----- | ---- | ----- | | | self | other | self | other | self | other | | 0 | NaN | NaN | NaN | NaN | NaN | NaN | | 1 | NaN | NaN | 22 | 32 | NaN | NaN | | 2 | NaN | NaN | 32 | 22 | 33 | 44 | Preserving the original values from both DataFrames can be done by both parameters: ```python df.compare(df2, keep_shape=True, keep_equal=True) ``` The result is: | | col\_1 | col\_2 | col\_3 | | | | | - | ------ | ------ | ------ | ----- | ---- | ----- | | | self | other | self | other | self | other | | 0 | 11 | 11 | 12 | 12 | 13 | 13 | | 1 | 21 | 21 | 22 | 32 | 23 | 23 | | 2 | 31 | 31 | 32 | 22 | 33 | 44 | ## 5\. Compare Two Pandas DataFrames Highlighting Differences Finally let's cover how to highlight the difference between both DataFrames. Again we are going to use the method `compare()` with a combination of styling. ### 5.1\. Compare and Highlight Difference between Two DataFrames To compare two DataFrames get the difference and highlight them use the code below: ```python df_mask = df.compare(df2, keep_shape=True).notnull().astype('int') df_compare = df.compare(df2, keep_shape=True, keep_equal=True) def apply_color(x): colors = {1: 'lightblue', 0: 'white'} return df_mask.applymap(lambda val: 'background-color: {}'.format(colors.get(val,''))) df_compare.style.apply(apply_color, axis=None) ``` ### 5.2\. Explanation of the steps First we are going to build a mask for our styling: ```python df_diff = df.compare(df2, keep_shape=True) df_mask = df_diff.notnull().astype('int') ``` The mask will look like: | | col\_1 | col\_2 | col\_3 | | | | | - | ------ | ------ | ------ | ----- | ---- | ----- | | | self | other | self | other | self | other | | 0 | 0 | 0 | 0 | 0 | 0 | 0 | | 1 | 0 | 0 | 1 | 1 | 0 | 0 | | 2 | 0 | 0 | 1 | 1 | 1 | 1 | Next is to build comparison DataFrame: ```python df_compare = df.compare(df2, keep_shape=True, keep_equal=True) df_compare ``` the same as the previous step. The last thing to do is apply styling based on the mask to the comparison DataFrame: ```python def apply_color(x): colors = {1: 'lightblue', 0: 'white'} return df_mask.applymap(lambda val: 'background-color: {}'.format(colors.get(val,''))) df_compare.style.apply(apply_color, axis=None) ``` the final result is: | | col\_1 | col\_2 | col\_3 | | | | | - | ------ | ------ | ------ | ----- | ---- | ----- | | | self | other | self | other | self | other | | 0 | 11 | 11 | 12 | 12 | 13 | 13 | | 1 | 21 | 21 | 22 | 32 | 23 | 23 | | 2 | 31 | 31 | 32 | 22 | 33 | 44 | ## 6\. Conclusion To sum up, we looked at how to **compare two DataFrames in Pandas side by side.** We focused on finding differences, keeping the same shape and the equal values. Finally we saw how to **highlight the differences while showing DataFrames side by side**. ### How to Get Today's Date in Pandas URL: https://datascientyst.com/get-todays-date-pandas/ Last updated: 2022-02-15T11:30:17.000Z ## 1\. Overview In this quick guide, we're going to look at **today's date in Pandas**. How to get **today with or without time, how to subtract days or date from it**. ## 2\. Get today's date in Pandas To **get today's date as datetime in Pandas** we can use the method `to_datetime()` and pass parameter - `today`. Below we are creating new column with Timestamp of today: ```python pd.to_datetime("today") df['today'] = pd.to_datetime("today") ``` The result is Timestamp of today's date and new column with the same date: ``` Timestamp('2022-02-15 12:00:42.371569') ``` ## 3\. Pandas get today's date without time Depending on the final data type there are several options how to extract dates in Pandas: - `string` \- `strftime()` - `datetime` \- method `date()` ### 3.1 Get Today's Date in Pandas - `strftime()` and format %m/%d/%Y If you like to **get today as a string** in custom date format then we can use method - `strftime()`: ```python pd.to_datetime("today").strftime("%m/%d/%Y") ``` the result is a string with the current date: ``` '02/15/2022' ``` ### 3.2 Get Today's Date in Pandas - method `date()` Pandas offer method `date()` which returns datetime from a given date. For today we can chain the methods in order to get today's date: ```python pd.to_datetime("today").date() ``` The result is: ``` datetime.date(2022, 2, 15) ``` ## 4\. Today date minus days or date in Pandas There are two main scenarious for dates and subtraction: - date minus days with result in new date - date minus another date - result in timedelta ### 4.1 Today date minus N days To s**ubtract days from todays date in Pandas** we can use the following format - `n * obj.freq`. We need to provide two parameters: - n - number of days. In this example is 10 - unit - the frequency of the object - D for days ```python pd.to_datetime("today") - pd.Timedelta(10, unit='D') ``` The result is a date with difference 10 days from today: ``` Timestamp('2022-02-05 13:23:10.382430') ``` **Note:** The subtracting from integer or float is deprecated and will raise error: > TypeError: Addition/subtraction of integers and integer-arrays with Timestamp is no longer supported. Instead of adding/subtracting `n`, use `n * obj.freq` ### 4.2 Today date minus another date In pandas is possible to perform subtraction of two dates with minus operator: ```python td = pd.to_datetime("today") - pd.to_datetime("02/10/2022") ``` The result will be timedelta: ``` Timedelta('5 days 12:47:35.699206') ``` Days can be extracted by using attribute - `days` from the result- `td.days` \- 5. ## 5\. Conclusion In this article, we looked at different solutions for getting today's date in Pandas. We covered most popular questions related to the dates and today in Pandas: remove time, subtract days or date. ### How to Add Gradient Color to Date Column in Pandas URL: https://datascientyst.com/add-gradient-color-date-column-pandas/ Last updated: 2022-02-04T15:55:44.000Z ## 1\. Overview In this short tutorial, we are going **to add gradient color visualization for dates in Pandas DataFrame.** This is helpful when you need to distinguish dates based on their values. ## 2\. Setup We will use simple DataFrame: ```python import numpy as np import pandas as pd from pandas.util.testing import makeTimeSeries s = makeTimeSeries(5) cols = ['col_1', 'col_2'] df = pd.DataFrame(abs(np.random.randn(5, 2)), columns=cols) df['date'] = s.index ``` data: | | col\_1 | col\_2 | date | | - | -------- | -------- | ---------- | | 0 | 0.182079 | 0.203433 | 2000-01-03 | | 1 | 0.037811 | 0.096004 | 2000-01-04 | | 2 | 0.479913 | 0.802299 | 2000-01-05 | | 3 | 1.881718 | 1.605743 | 2000-01-06 | | 4 | 0.319824 | 0.912157 | 2000-01-07 | ## 3\. Steps to Add gradient to Date Column in Pandas ### 3.1\. Add new column days since We will start by adding a new column which will have the difference between some date or today and the dates in the column: ```python from datetime import datetime df['days since'] = (df['date'] - datetime.today()).dt.days ``` DataFrame will looks like: | | col\_1 | col\_2 | date | days since | | - | -------- | -------- | ------------------- | ---------- | | 0 | 0.754816 | 0.107682 | 2000-01-03 00:00:00 | \-8069 | | 1 | 0.656553 | 0.042157 | 2000-01-04 00:00:00 | \-8068 | | 2 | 0.953792 | 1.593115 | 2000-01-05 00:00:00 | \-8067 | | 3 | 0.671552 | 1.810286 | 2000-01-06 00:00:00 | \-8066 | | 4 | 0.159057 | 0.548218 | 2000-01-07 00:00:00 | \-8065 | ### 3.2\. Add gradient color to days column as a Styler Then we are going to use method `background_gradient` to add gradient color to the whole DataFrame: ```python style1 = df.style.background_gradient(cmap='Greens') ``` ### 3.3\. Create new DataFrame and reuse the style Next we are going do the following steps: - copy the first DataFrame as `df2` - create new Styler - hide column 'date' - change column 'days since' to 'date' column - reuse the first Styler ```python style1 = df.style.background_gradient(cmap='Greens') df2 = df style2 = df2.style.hide(['date'], axis='columns') df2['days since'] = df2['date'] style2.use(style1.export()) ``` Dates after this step will be shown as: `2000-01-03 00:00:00З` ### 3.4\. Change date format in the Pandas style object Finally we are going to change the date format in order to make ir more readable by adding `.format({"days since": lambda t: t.strftime("%m/%d/%Y")})` The dates now will be displayed as: `01/03/2000` We can use any format thanks to `strftime` ## 4\. Add gradient to Date Column in Pandas - Full code Finally we can see the full code and the final result below: ```python style1 = df.style.background_gradient(cmap='Greens') df2 = df style2 = df2.style.hide(['date'], axis='columns') df2['days since'] = df2['date'] style2.use(style1.export()).format({"days since": lambda t: t.strftime("%m/%d/%Y")}) ``` result: | | col\_1 | col\_2 | days since | | - | -------- | -------- | ---------- | | 0 | 0.754816 | 0.107682 | 01/03/2000 | | 1 | 0.656553 | 0.042157 | 01/04/2000 | | 2 | 0.953792 | 1.593115 | 01/05/2000 | | 3 | 0.671552 | 1.810286 | 01/06/2000 | | 4 | 0.159057 | 0.548218 | 01/07/2000 | ## 5\. Conclusion We saw how to add gradient color for dates. We cover all the steps with detailed explanation. We also learned how to copy styles between two DataFrames and how to change the format of dates. ### How to Display Pandas DataFrame As a Heatmap URL: https://datascientyst.com/display-pandas-dataframe-heatmap/ Last updated: 2022-02-04T15:55:05.000Z ## 1\. Overview In this tutorial, we'll learn how to display Pandas DataFrame as a heatmap. So we might start with: what is a heatmap in Data Science? According to wikipedia: > A heat map (or heatmap) is a data visualization technique that shows the magnitude of a phenomenon as color in two dimensions. ## 2\. Setup We are going to create test DataFrame following two articles: - [How to Create a Pandas DataFrame of Random Integers ](https://datascientyst.com/how-to-create-a-dataframe-of-random-integers-with-pandas/) - [How to Easily Create Dummy DataFrame with Test Data?](https://datascientyst.com/create-easily-dummy-dataframe-test-data/#7-dummy-dataframe-with-mixed-data) We are using `np.random.randn(10, 3)` to create DataFrame with 10 rows and 3 columns- with random values: ```python import numpy as np import pandas as pd from pandas.util.testing import makePeriodSeries s = makeTimeSeries(10) cols = ['col_1', 'col_2'] df = pd.DataFrame(abs(np.random.randn(10, 2)), columns=cols) df['item'] = 'item: ' + df.index.astype(str) df['date'] = s.index ``` | | col\_1 | col\_2 | item | date | | - | -------- | -------- | ------- | ---------- | | 0 | 0.448082 | 0.334594 | item: 0 | 2000-01-03 | | 1 | 0.727165 | 0.349513 | item: 1 | 2000-01-04 | | 2 | 0.628442 | 0.485067 | item: 2 | 2000-01-05 | | 3 | 0.193080 | 1.361732 | item: 3 | 2000-01-06 | | 4 | 0.358394 | 0.746719 | item: 4 | 2000-01-07 | | 5 | 0.089303 | 1.600171 | item: 5 | 2000-01-10 | | 6 | 0.126041 | 0.943686 | item: 6 | 2000-01-11 | | 7 | 0.002382 | 0.516401 | item: 7 | 2000-01-12 | | 8 | 0.058525 | 1.233783 | item: 8 | 2000-01-13 | | 9 | 1.433061 | 1.703305 | item: 9 | 2000-01-14 | ## 3\. Pandas: Display DataFrame as heatmap with style.background\_gradient Pandas offer method `style.background_gradient()` which helps us very easily to create beautiful colored heatmap: ```python df.style.background_gradient(cmap='Greens') ``` The background gradient it will applied only for the numeric columns: | | col\_1 | col\_2 | item | date | | - | -------- | -------- | ------- | ---------- | | 0 | 0.448082 | 0.334594 | item: 0 | 2000-01-03 | | 1 | 0.727165 | 0.349513 | item: 1 | 2000-01-04 | | 2 | 0.628442 | 0.485067 | item: 2 | 2000-01-05 | | 3 | 0.193080 | 1.361732 | item: 3 | 2000-01-06 | | 4 | 0.358394 | 0.746719 | item: 4 | 2000-01-07 | | 5 | 0.089303 | 1.600171 | item: 5 | 2000-01-10 | | 6 | 0.126041 | 0.943686 | item: 6 | 2000-01-11 | | 7 | 0.002382 | 0.516401 | item: 7 | 2000-01-12 | | 8 | 0.058525 | 1.233783 | item: 8 | 2000-01-13 | | 9 | 1.433061 | 1.703305 | item: 9 | 2000-01-14 | The method `background_gradient()` take as argument `cmap` which can have different values like: - `Blues` - `Greens` To learn more about Pandas colors and palettes please visit: - [How to Get a List of N Different Colors and Names in Python/Pandas ](https://datascientyst.com/get-list-of-n-different-colors-names-python-pandas/) - [Choosing Colormaps in Matplotlib](https://matplotlib.org/stable/tutorials/colors/colormaps.html?ref=datascientyst.com) ## 4\. Seaborn: Display DataFrame as heatmap with sns.heatmap There is a library for data visualization called [Seaborn: statistical data visualization](https://seaborn.pydata.org/?ref=datascientyst.com). This library offers method called: [seaborn.heatmap()](https://seaborn.pydata.org/generated/seaborn.heatmap.html?ref=datascientyst.com) The method works only on numerical values. So we can use it as follow: ```python import seaborn as sns sns.heatmap(df[['col_1', 'col_2']]) ``` the DataFrame will looks like: ![display-pandas-dataframe-heatmap-seaborn-sns](https://datascientyst.com/content/images/2022/02/display-pandas-dataframe-heatmap-seaborn-sns.png) If you try to call the method: `sns.heatmap()` on the whole DataFrame you will get an error: > ValueError: could not convert string to float: 'item: 0' Another way to solve the error is by pivoting data on some columns. For example we can pivot on columns: - "date" - "item" and get as values column - "col\_1": ```python sns.heatmap(df.pivot("date", "item",values='col_1')) ``` This will convert the DataFrame into beautiful heatmap: ![display-pandas-dataframe-heatmap-seaborn-pivot](https://datascientyst.com/content/images/2022/02/display-pandas-dataframe-heatmap-seaborn-pivot.png) Again we can provide parameter `cmap` which can take similar values as the `background_gradient()`. ## 5\. Interactive heatmap with Plotly If you like to make your DataFrame as aa interactive heatmap then you can use library called: - [Plotly: The front end for ML and data science models](https://plotly.com/?ref=datascientyst.com) Again as Seaborn we need to use only numeric values: ```python import plotly.express as px fig = px.imshow(df[['col_1', 'col_2']]) fig.show() ``` Otherwise errors will be raised. The resulted heatmap will looks like: ![display-pandas-dataframe-interactive-heatmap-plotly](https://datascientyst.com/content/images/2022/02/display-pandas-dataframe-interactive-heatmap-plotly.png) For categorical data we can use `pivot()` or similar operation in order to make it good for plotting as a heatmap. The error is a bit different: > TypeError: Object of type Period is not JSON serializable ## 6\. Conclusion We covered the most popular ways to convert DataFrame to: - heatmap for numeric and non numeric data - heatmap with seaborn - data transformation for categorical data with pivot - interactive heatmap This article will help you to select the best way to present your numeric data. ### How to Convert First Row to Header Column in Pandas DataFrame URL: https://datascientyst.com/convert-first-row-header-column-pandas-dataframe/ Last updated: 2022-02-04T16:09:20.000Z ## 1\. Overview This guide describes how to convert first or other rows as a header in Pandas DataFrame. We will cover several different examples with details. If you are using read\_csv() method you can learn more - about headers: [How to Read Excel or CSV With Multiple Line Headers Using Pandas ](https://datascientyst.com/read-excel-csv-multiple-line-headers-using-pandas/) - [How to Reset Column Names (Index) in Pandas](https://datascientyst.com/reset-column-names-index-pandas/) ## 2\. Setup We are going to work with simple DataFrame created by: ```python import pandas as pd df = pd.DataFrame([('head_1', 'head_2', 'head_3' ), ('val_11','val_12','val_13'), ('val_21','val_22','val_23'), ('val_31','val_32','val_33')]) ``` Final DataFrame looks like: | | 0 | 1 | 2 | | - | ------- | ------- | ------- | | 0 | head\_1 | head\_2 | head\_3 | | 1 | val\_11 | val\_12 | val\_13 | | 2 | val\_21 | val\_22 | val\_23 | | 3 | val\_31 | val\_32 | val\_33 | From this DataFrame we can conclude that the first row of it should be used as a header. ## 3\. Using First Row as a Header with df.rename() The first solution is to combine two Pandas methods: - [pandas.DataFrame.rename](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.rename.html?ref=datascientyst.com) - [pandas.DataFrame.drop](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.drop.html?ref=datascientyst.com) The method `.rename(columns=)` expects to be iterable with the column names. To select the first row we are going to use `iloc` \- `df.iloc[0]`. Finally we need to drop the first row which was used as a header by `drop(df.index[0])`: ```python df.rename(columns=df.iloc[0]).drop(df.index[0]) ``` The result is: | | head\_1 | head\_2 | head\_3 | | - | ------- | ------- | ------- | | 1 | val\_11 | val\_12 | val\_13 | | 2 | val\_21 | val\_22 | val\_23 | | 3 | val\_31 | val\_32 | val\_33 | For other rows we can change the index - 0\. For example to use the last row as header: -1 - `df.iloc[-1]`. To make the change permanent we need to use `inplace = True` or reassign the DataFrame. ## 4\. Using First Row as a Header with pd.DataFrame() Another solution is to create new DataFrame by using the values from the first one - up to the first row: `df.values[1:]` Use the column header from the first row of the existing DataFrame. ```python pd.DataFrame(df.values[1:], columns=df.iloc[0]) ``` The result is exactly the same as the previous solution. There are several differences: - This solution might be slower for bigger DataFrames - It creates new DataFrame - It may change the dtypes of the new DataFrame ## 5\. Conclusion In this short post we saw how to use a row as a header in Pandas. We covered also several Pandas methods like: `iloc()`, `rename()` and `drop()` ### Full List of Named Colors in Pandas and Python URL: https://datascientyst.com/full-list-named-colors-pandas-python-matplotlib/ Last updated: 2023-03-31T06:59:33.000Z ## 1\. Overview This article is a reference of all named colors in Pandas. It **shows a list of more than 1200+ named colors in Python, Matplotlib and Pandas.** They are based on the Python library Matplotlib. The work in based on two articles: - [How to Get a List of N Different Colors and Names in Python/Pandas ](https://datascientyst.com/get-list-of-n-different-colors-names-python-pandas/) - [List of named colors - Matplotlib](https://matplotlib.org/stable/gallery/color/named%5Fcolors.html?ref=datascientyst.com) We are going to build different color palettes and different conversion techniques. The image below shows some of the colors: ![full-list-named-colors-pandas-python-matplotlib](https://datascientyst.com/content/images/2022/10/full-list-named-colors-pandas-python-matplotlib.webp) ## 2\. Convert colors in Matplotlib, Python and Pandas To learn more about the color conversion in Python you can check 5th step of the above article: [Working with color names and color values](https://datascientyst.com/get-list-of-n-different-colors-names-python-pandas/#step-5-working-with-color-names-and-color-values-format-html-and-css-in-python) In this section we will cover the next conversions: - RGB to HEX = `(0.0, 1.0, 1.0)` \-> `#00FFFF` - HEX to RGB = `#00FFFF` \-> `(0.0, 1.0, 1.0)` ### 2.1\. Convert RGB to HEX color in Python and Pandas Let's start with conversion of RGB color in decimal code to HEX code. To do so we are going to use Matplotlib method: `mcolors.rgb2hex()`: ```python import matplotlib.colors as mcolors mcolors.rgb2hex((0.0, 1.0, 1.0)) ``` result: ``` '#00ffff' ``` ### 2.2\. Convert RGB to HSL in Pandas To convert from RGB to HSL in Pandas we are going to use method: `rgb_to_hsv`: ```python mcolors.rgb_to_hsv((0, 0, 1)) ``` the output is array of hue, saturation and lightness: ``` array([0.66666667, 1. , 1. ]) ``` ### 2.3\. Convert HEX to RGB format in Python and Pandas Similarly we can do the HEX to RGB conversion by using method - \`mcolors.hex2color(): ```python import matplotlib.colors as mcolors mcolors.hex2color('#40E0D0') ``` result: ``` (0.25098039215686274, 0.8784313725490196, 0.8156862745098039) ``` The result has high decimal precision. If you like to reduce the decimal numbers we can use list comprehension: ```python [round(c, 5) for c in (0.25098039215686274, 0.8784313725490196, 0.8156862745098039)] ``` will give us: ``` [0.25098, 0.87843, 0.81569] ``` To apply to a column in Pandas DataFrame we can use a lambda expression: ```python df_colors['rgb'] = df_colors['rgb'].apply(lambda x:[round(c, 5) for c in x]) ``` ### 2.4\. Convert RGB to HEX color for whole column in Pandas We can also apply the conversion to a single or multiple columns in Pandas DataFrame. For this purpose we are going to use method `apply()`: ```python df_colors['hex'] = df_colors['rgb'].apply(mcolors.rgb2hex) ``` The following example demonstrate this: ```python import pandas as pd colors = { 'name': mcolors.BASE_COLORS.keys(), 'rgb': mcolors.BASE_COLORS.values() } df_colors = pd.DataFrame(colors) df_colors['hex'] = df_colors['rgb'].apply(mcolors.rgb2hex) ``` The result DataFrame with converted RGB column to HEX one: | | name | rgb | hex | | - | ---- | --------------- | ------- | | 0 | b | (0, 0, 1) | #0000ff | | 1 | g | (0, 0.5, 0) | #008000 | | 2 | r | (1, 0, 0) | #ff0000 | | 3 | c | (0, 0.75, 0.75) | #00bfbf | | 4 | m | (0.75, 0, 0.75) | #bf00bf | ## 3\. BASE\_COLORS: Named Colors in Pandas We will start with the most basic colors with names from Matplotlib - `mcolors.BASE_COLORS`. Below you can find a table with the color name, RGB and HEX values plus display of the color: We are going to use the next code to generate the values: ```python import pandas as pd def format_color_groups(df, color): x = df.copy() i = 0 for factor in color: x.iloc[i, :-1] = '' style = f'background-color: {color[i]}' x.loc[i, 'display color as background'] = style i = i + 1 return x colors = { 'name': mcolors.BASE_COLORS.keys(), 'rgb': mcolors.BASE_COLORS.values() } df_colors = pd.DataFrame(colors) df_colors['hex'] = df_colors['rgb'].apply(mcolors.rgb2hex) df_colors['display color as background'] = '' df_colors.style.apply(format_color_groups, color=df_colors.hex, axis=None) ``` The list of the basic colors in Pandas is shown below: | | name | rgb | hex | display color as background | | - | ---- | --------------- | ------- | --------------------------- | | 0 | b | (0, 0, 1) | #0000ff | | | 1 | g | (0, 0.5, 0) | #008000 | | | 2 | r | (1, 0, 0) | #ff0000 | | | 3 | c | (0, 0.75, 0.75) | #00bfbf | | | 4 | m | (0.75, 0, 0.75) | #bf00bf | | | 5 | y | (0.75, 0.75, 0) | #bfbf00 | | | 6 | k | (0, 0, 0) | #000000 | | | 7 | w | (1, 1, 1) | #ffffff | | ## 4\. TABLEAU\_COLORS: List of Named Colors in Python and Pandas Next we will cover the list of TABLEAU\_COLORS in Pandas. They are limited to the most popular colors only. The code below will list all of them: ```python import pandas as pd def format_color_groups(df, color): x = df.copy() i = 0 for factor in color: x.iloc[i, :-1] = '' style = f'background-color: {color[i]}' x.loc[i, 'display color as background'] = style i = i + 1 return x colors = { 'name': mcolors.TABLEAU_COLORS.keys(), 'hex': mcolors.TABLEAU_COLORS.values() } df_colors = pd.DataFrame(colors) df_colors['rgb'] = df_colors['hex'].apply(mcolors.hex2color) df_colors['rgb'] = df_colors['rgb'].apply(lambda x:[round(c, 5) for c in x]) df_colors['display color as background'] = '' df_colors.style.apply(format_color_groups, color=df_colors.hex, axis=None) ``` And the table is: | | name | hex | rgb | display color as background | | - | ---------- | ------- | ----------------------------- | --------------------------- | | 0 | tab:blue | #1f77b4 | \[0.12157, 0.46667, 0.70588\] | | | 1 | tab:orange | #ff7f0e | \[1.0, 0.49804, 0.0549\] | | | 2 | tab:green | #2ca02c | \[0.17255, 0.62745, 0.17255\] | | | 3 | tab:red | #d62728 | \[0.83922, 0.15294, 0.15686\] | | | 4 | tab:purple | #9467bd | \[0.58039, 0.40392, 0.74118\] | | | 5 | tab:brown | #8c564b | \[0.54902, 0.33725, 0.29412\] | | | 6 | tab:pink | #e377c2 | \[0.8902, 0.46667, 0.76078\] | | | 7 | tab:gray | #7f7f7f | \[0.49804, 0.49804, 0.49804\] | | | 8 | tab:olive | #bcbd22 | \[0.73725, 0.74118, 0.13333\] | | | 9 | tab:cyan | #17becf | \[0.0902, 0.7451, 0.81176\] | | ## 5\. CSS4\_COLORS colors in Python and Pandas The CSS4\_COLORS contains many different named colors which can be used in Python and Matplotlib. We can list all of them by: ```python def format_color_groups(df, color): x = df.copy() i = 0 for factor in color: x.iloc[i, :-1] = '' style = f'background-color: {color[i]}' x.loc[i, 'display color as background'] = style i = i + 1 return x colors = { 'name': mcolors.CSS4_COLORS.keys(), 'hex': mcolors.CSS4_COLORS.values() } df_colors = pd.DataFrame(colors) df_colors['rgb'] = df_colors['hex'].apply(mcolors.hex2color) df_colors['rgb'] = df_colors['rgb'].apply(lambda x:[round(c, 5) for c in x]) df_colors['display color as background'] = '' df_colors.style.apply(format_color_groups, color=df_colors.hex, axis=None) ``` We can find all the colors from this pallet below: | | name | hex | rgb | display color as background | | --- | -------------------- | ------- | ----------------------------- | --------------------------- | | 0 | aliceblue | #F0F8FF | \[0.94118, 0.97255, 1.0\] | | | 1 | antiquewhite | #FAEBD7 | \[0.98039, 0.92157, 0.84314\] | | | 2 | aqua | #00FFFF | \[0.0, 1.0, 1.0\] | | | 3 | aquamarine | #7FFFD4 | \[0.49804, 1.0, 0.83137\] | | | 4 | azure | #F0FFFF | \[0.94118, 1.0, 1.0\] | | | 5 | beige | #F5F5DC | \[0.96078, 0.96078, 0.86275\] | | | 6 | bisque | #FFE4C4 | \[1.0, 0.89412, 0.76863\] | | | 7 | black | #000000 | \[0.0, 0.0, 0.0\] | | | 8 | blanchedalmond | #FFEBCD | \[1.0, 0.92157, 0.80392\] | | | 9 | blue | #0000FF | \[0.0, 0.0, 1.0\] | | | 10 | blueviolet | #8A2BE2 | \[0.54118, 0.16863, 0.88627\] | | | 11 | brown | #A52A2A | \[0.64706, 0.16471, 0.16471\] | | | 12 | burlywood | #DEB887 | \[0.87059, 0.72157, 0.52941\] | | | 13 | cadetblue | #5F9EA0 | \[0.37255, 0.61961, 0.62745\] | | | 14 | chartreuse | #7FFF00 | \[0.49804, 1.0, 0.0\] | | | 15 | chocolate | #D2691E | \[0.82353, 0.41176, 0.11765\] | | | 16 | coral | #FF7F50 | \[1.0, 0.49804, 0.31373\] | | | 17 | cornflowerblue | #6495ED | \[0.39216, 0.58431, 0.92941\] | | | 18 | cornsilk | #FFF8DC | \[1.0, 0.97255, 0.86275\] | | | 19 | crimson | #DC143C | \[0.86275, 0.07843, 0.23529\] | | | 20 | cyan | #00FFFF | \[0.0, 1.0, 1.0\] | | | 21 | darkblue | #00008B | \[0.0, 0.0, 0.5451\] | | | 22 | darkcyan | #008B8B | \[0.0, 0.5451, 0.5451\] | | | 23 | darkgoldenrod | #B8860B | \[0.72157, 0.52549, 0.04314\] | | | 24 | darkgray | #A9A9A9 | \[0.66275, 0.66275, 0.66275\] | | | 25 | darkgreen | #006400 | \[0.0, 0.39216, 0.0\] | | | 26 | darkgrey | #A9A9A9 | \[0.66275, 0.66275, 0.66275\] | | | 27 | darkkhaki | #BDB76B | \[0.74118, 0.71765, 0.41961\] | | | 28 | darkmagenta | #8B008B | \[0.5451, 0.0, 0.5451\] | | | 29 | darkolivegreen | #556B2F | \[0.33333, 0.41961, 0.18431\] | | | 30 | darkorange | #FF8C00 | \[1.0, 0.54902, 0.0\] | | | 31 | darkorchid | #9932CC | \[0.6, 0.19608, 0.8\] | | | 32 | darkred | #8B0000 | \[0.5451, 0.0, 0.0\] | | | 33 | darksalmon | #E9967A | \[0.91373, 0.58824, 0.47843\] | | | 34 | darkseagreen | #8FBC8F | \[0.56078, 0.73725, 0.56078\] | | | 35 | darkslateblue | #483D8B | \[0.28235, 0.23922, 0.5451\] | | | 36 | darkslategray | #2F4F4F | \[0.18431, 0.3098, 0.3098\] | | | 37 | darkslategrey | #2F4F4F | \[0.18431, 0.3098, 0.3098\] | | | 38 | darkturquoise | #00CED1 | \[0.0, 0.80784, 0.81961\] | | | 39 | darkviolet | #9400D3 | \[0.58039, 0.0, 0.82745\] | | | 40 | deeppink | #FF1493 | \[1.0, 0.07843, 0.57647\] | | | 41 | deepskyblue | #00BFFF | \[0.0, 0.74902, 1.0\] | | | 42 | dimgray | #696969 | \[0.41176, 0.41176, 0.41176\] | | | 43 | dimgrey | #696969 | \[0.41176, 0.41176, 0.41176\] | | | 44 | dodgerblue | #1E90FF | \[0.11765, 0.56471, 1.0\] | | | 45 | firebrick | #B22222 | \[0.69804, 0.13333, 0.13333\] | | | 46 | floralwhite | #FFFAF0 | \[1.0, 0.98039, 0.94118\] | | | 47 | forestgreen | #228B22 | \[0.13333, 0.5451, 0.13333\] | | | 48 | fuchsia | #FF00FF | \[1.0, 0.0, 1.0\] | | | 49 | gainsboro | #DCDCDC | \[0.86275, 0.86275, 0.86275\] | | | 50 | ghostwhite | #F8F8FF | \[0.97255, 0.97255, 1.0\] | | | 51 | gold | #FFD700 | \[1.0, 0.84314, 0.0\] | | | 52 | goldenrod | #DAA520 | \[0.8549, 0.64706, 0.12549\] | | | 53 | gray | #808080 | \[0.50196, 0.50196, 0.50196\] | | | 54 | green | #008000 | \[0.0, 0.50196, 0.0\] | | | 55 | greenyellow | #ADFF2F | \[0.67843, 1.0, 0.18431\] | | | 56 | grey | #808080 | \[0.50196, 0.50196, 0.50196\] | | | 57 | honeydew | #F0FFF0 | \[0.94118, 1.0, 0.94118\] | | | 58 | hotpink | #FF69B4 | \[1.0, 0.41176, 0.70588\] | | | 59 | indianred | #CD5C5C | \[0.80392, 0.36078, 0.36078\] | | | 60 | indigo | #4B0082 | \[0.29412, 0.0, 0.5098\] | | | 61 | ivory | #FFFFF0 | \[1.0, 1.0, 0.94118\] | | | 62 | khaki | #F0E68C | \[0.94118, 0.90196, 0.54902\] | | | 63 | lavender | #E6E6FA | \[0.90196, 0.90196, 0.98039\] | | | 64 | lavenderblush | #FFF0F5 | \[1.0, 0.94118, 0.96078\] | | | 65 | lawngreen | #7CFC00 | \[0.48627, 0.98824, 0.0\] | | | 66 | lemonchiffon | #FFFACD | \[1.0, 0.98039, 0.80392\] | | | 67 | lightblue | #ADD8E6 | \[0.67843, 0.84706, 0.90196\] | | | 68 | lightcoral | #F08080 | \[0.94118, 0.50196, 0.50196\] | | | 69 | lightcyan | #E0FFFF | \[0.87843, 1.0, 1.0\] | | | 70 | lightgoldenrodyellow | #FAFAD2 | \[0.98039, 0.98039, 0.82353\] | | | 71 | lightgray | #D3D3D3 | \[0.82745, 0.82745, 0.82745\] | | | 72 | lightgreen | #90EE90 | \[0.56471, 0.93333, 0.56471\] | | | 73 | lightgrey | #D3D3D3 | \[0.82745, 0.82745, 0.82745\] | | | 74 | lightpink | #FFB6C1 | \[1.0, 0.71373, 0.75686\] | | | 75 | lightsalmon | #FFA07A | \[1.0, 0.62745, 0.47843\] | | | 76 | lightseagreen | #20B2AA | \[0.12549, 0.69804, 0.66667\] | | | 77 | lightskyblue | #87CEFA | \[0.52941, 0.80784, 0.98039\] | | | 78 | lightslategray | #778899 | \[0.46667, 0.53333, 0.6\] | | | 79 | lightslategrey | #778899 | \[0.46667, 0.53333, 0.6\] | | | 80 | lightsteelblue | #B0C4DE | \[0.6902, 0.76863, 0.87059\] | | | 81 | lightyellow | #FFFFE0 | \[1.0, 1.0, 0.87843\] | | | 82 | lime | #00FF00 | \[0.0, 1.0, 0.0\] | | | 83 | limegreen | #32CD32 | \[0.19608, 0.80392, 0.19608\] | | | 84 | linen | #FAF0E6 | \[0.98039, 0.94118, 0.90196\] | | | 85 | magenta | #FF00FF | \[1.0, 0.0, 1.0\] | | | 86 | maroon | #800000 | \[0.50196, 0.0, 0.0\] | | | 87 | mediumaquamarine | #66CDAA | \[0.4, 0.80392, 0.66667\] | | | 88 | mediumblue | #0000CD | \[0.0, 0.0, 0.80392\] | | | 89 | mediumorchid | #BA55D3 | \[0.72941, 0.33333, 0.82745\] | | | 90 | mediumpurple | #9370DB | \[0.57647, 0.43922, 0.85882\] | | | 91 | mediumseagreen | #3CB371 | \[0.23529, 0.70196, 0.44314\] | | | 92 | mediumslateblue | #7B68EE | \[0.48235, 0.40784, 0.93333\] | | | 93 | mediumspringgreen | #00FA9A | \[0.0, 0.98039, 0.60392\] | | | 94 | mediumturquoise | #48D1CC | \[0.28235, 0.81961, 0.8\] | | | 95 | mediumvioletred | #C71585 | \[0.78039, 0.08235, 0.52157\] | | | 96 | midnightblue | #191970 | \[0.09804, 0.09804, 0.43922\] | | | 97 | mintcream | #F5FFFA | \[0.96078, 1.0, 0.98039\] | | | 98 | mistyrose | #FFE4E1 | \[1.0, 0.89412, 0.88235\] | | | 99 | moccasin | #FFE4B5 | \[1.0, 0.89412, 0.7098\] | | | 100 | navajowhite | #FFDEAD | \[1.0, 0.87059, 0.67843\] | | | 101 | navy | #000080 | \[0.0, 0.0, 0.50196\] | | | 102 | oldlace | #FDF5E6 | \[0.99216, 0.96078, 0.90196\] | | | 103 | olive | #808000 | \[0.50196, 0.50196, 0.0\] | | | 104 | olivedrab | #6B8E23 | \[0.41961, 0.55686, 0.13725\] | | | 105 | orange | #FFA500 | \[1.0, 0.64706, 0.0\] | | | 106 | orangered | #FF4500 | \[1.0, 0.27059, 0.0\] | | | 107 | orchid | #DA70D6 | \[0.8549, 0.43922, 0.83922\] | | | 108 | palegoldenrod | #EEE8AA | \[0.93333, 0.9098, 0.66667\] | | | 109 | palegreen | #98FB98 | \[0.59608, 0.98431, 0.59608\] | | | 110 | paleturquoise | #AFEEEE | \[0.68627, 0.93333, 0.93333\] | | | 111 | palevioletred | #DB7093 | \[0.85882, 0.43922, 0.57647\] | | | 112 | papayawhip | #FFEFD5 | \[1.0, 0.93725, 0.83529\] | | | 113 | peachpuff | #FFDAB9 | \[1.0, 0.8549, 0.72549\] | | | 114 | peru | #CD853F | \[0.80392, 0.52157, 0.24706\] | | | 115 | pink | #FFC0CB | \[1.0, 0.75294, 0.79608\] | | | 116 | plum | #DDA0DD | \[0.86667, 0.62745, 0.86667\] | | | 117 | powderblue | #B0E0E6 | \[0.6902, 0.87843, 0.90196\] | | | 118 | purple | #800080 | \[0.50196, 0.0, 0.50196\] | | | 119 | rebeccapurple | #663399 | \[0.4, 0.2, 0.6\] | | | 120 | red | #FF0000 | \[1.0, 0.0, 0.0\] | | | 121 | rosybrown | #BC8F8F | \[0.73725, 0.56078, 0.56078\] | | | 122 | royalblue | #4169E1 | \[0.2549, 0.41176, 0.88235\] | | | 123 | saddlebrown | #8B4513 | \[0.5451, 0.27059, 0.07451\] | | | 124 | salmon | #FA8072 | \[0.98039, 0.50196, 0.44706\] | | | 125 | sandybrown | #F4A460 | \[0.95686, 0.64314, 0.37647\] | | | 126 | seagreen | #2E8B57 | \[0.18039, 0.5451, 0.34118\] | | | 127 | seashell | #FFF5EE | \[1.0, 0.96078, 0.93333\] | | | 128 | sienna | #A0522D | \[0.62745, 0.32157, 0.17647\] | | | 129 | silver | #C0C0C0 | \[0.75294, 0.75294, 0.75294\] | | | 130 | skyblue | #87CEEB | \[0.52941, 0.80784, 0.92157\] | | | 131 | slateblue | #6A5ACD | \[0.41569, 0.35294, 0.80392\] | | | 132 | slategray | #708090 | \[0.43922, 0.50196, 0.56471\] | | | 133 | slategrey | #708090 | \[0.43922, 0.50196, 0.56471\] | | | 134 | snow | #FFFAFA | \[1.0, 0.98039, 0.98039\] | | | 135 | springgreen | #00FF7F | \[0.0, 1.0, 0.49804\] | | | 136 | steelblue | #4682B4 | \[0.27451, 0.5098, 0.70588\] | | | 137 | tan | #D2B48C | \[0.82353, 0.70588, 0.54902\] | | | 138 | teal | #008080 | \[0.0, 0.50196, 0.50196\] | | | 139 | thistle | #D8BFD8 | \[0.84706, 0.74902, 0.84706\] | | | 140 | tomato | #FF6347 | \[1.0, 0.38824, 0.27843\] | | | 141 | turquoise | #40E0D0 | \[0.25098, 0.87843, 0.81569\] | | | 142 | violet | #EE82EE | \[0.93333, 0.5098, 0.93333\] | | | 143 | wheat | #F5DEB3 | \[0.96078, 0.87059, 0.70196\] | | | 144 | white | #FFFFFF | \[1.0, 1.0, 1.0\] | | | 145 | whitesmoke | #F5F5F5 | \[0.96078, 0.96078, 0.96078\] | | | 146 | yellow | #FFFF00 | \[1.0, 1.0, 0.0\] | | | 147 | yellowgreen | #9ACD32 | \[0.60392, 0.80392, 0.19608\] | | ## 6\. XKCD\_Colors: colors in Pandas and Python Finally we will cover huge list of named colors in Python: XKCD\_Colors. The list contains about 950 colors. So we are going to list only some of them: The code for listing all 900+ XKCD\_Colors in Pandas is similar to the previous ones: ```python def format_color_groups(df, color): x = df.copy() i = 0 for factor in color: x.iloc[i, :-1] = '' style = f'background-color: {color[i]}' x.loc[i, 'display color as background'] = style i = i + 1 return x colors = { 'name': mcolors.XKCD_COLORS.keys(), 'hex': mcolors.XKCD_COLORS.values() } df_colors = pd.DataFrame(colors) df_colors['rgb'] = df_colors['hex'].apply(mcolors.hex2color) df_colors['rgb'] = df_colors['rgb'].apply(lambda x:[round(c, 5) for c in x]) df_colors['display color as background'] = '' df_colors.style.apply(format_color_groups, color=df_colors.hex, axis=None) ``` Sample of the first 15 colors: | | name | hex | rgb | display color as background | | -- | ---------------------- | ------- | ----------------------------- | --------------------------- | | 0 | xkcd:cloudy blue | #acc2d9 | \[0.67451, 0.76078, 0.85098\] | | | 1 | xkcd:dark pastel green | #56ae57 | \[0.33725, 0.68235, 0.34118\] | | | 2 | xkcd:dust | #b2996e | \[0.69804, 0.6, 0.43137\] | | | 3 | xkcd:electric lime | #a8ff04 | \[0.65882, 1.0, 0.01569\] | | | 4 | xkcd:fresh green | #69d84f | \[0.41176, 0.84706, 0.3098\] | | | 5 | xkcd:light eggplant | #894585 | \[0.53725, 0.27059, 0.52157\] | | | 6 | xkcd:nasty green | #70b23f | \[0.43922, 0.69804, 0.24706\] | | | 7 | xkcd:really light blue | #d4ffff | \[0.83137, 1.0, 1.0\] | | | 8 | xkcd:tea | #65ab7c | \[0.39608, 0.67059, 0.48627\] | | | 9 | xkcd:warm purple | #952e8f | \[0.58431, 0.18039, 0.56078\] | | | 10 | xkcd:yellowish tan | #fcfc81 | \[0.98824, 0.98824, 0.50588\] | | | 11 | xkcd:cement | #a5a391 | \[0.64706, 0.63922, 0.56863\] | | | 12 | xkcd:dark grass green | #388004 | \[0.21961, 0.50196, 0.01569\] | | | 13 | xkcd:dusty teal | #4c9085 | \[0.29804, 0.56471, 0.52157\] | | | 14 | xkcd:grey teal | #5e9b8a | \[0.36863, 0.60784, 0.54118\] | | ## 7\. Matplotlib colors palettes - cmaps In this section we can find list and view of all Matplotlib colors palettes. They are known as `cmap` and can be found in Pandas functions like parameters. In total there are 166 Matplotlib colors palettes: > \['magma', 'inferno', 'plasma', 'viridis', 'cividis', 'twilight', 'twilight\_shifted', 'turbo', 'Blues', 'BrBG', 'BuGn', 'BuPu', 'CMRmap', 'GnBu', 'Greens', 'Greys', 'OrRd', 'Oranges', 'PRGn', 'PiYG', 'PuBu', 'PuBuGn', 'PuOr', 'PuRd', 'Purples', 'RdBu', 'RdGy', 'RdPu', 'RdYlBu', 'RdYlGn', 'Reds', 'Spectral', 'Wistia', 'YlGn', 'YlGnBu', 'YlOrBr', 'YlOrRd', 'afmhot', 'autumn', 'binary', 'bone', 'brg', 'bwr', 'cool', 'coolwarm', 'copper', 'cubehelix', 'flag', 'gist\_earth', 'gist\_gray', 'gist\_heat', 'gist\_ncar', 'gist\_rainbow', 'gist\_stern', 'gist\_yarg', 'gnuplot', 'gnuplot2', 'gray', 'hot', 'hsv', 'jet', 'nipy\_spectral', 'ocean', 'pink', 'prism', 'rainbow', 'seismic', 'spring', 'summer', 'terrain', 'winter', 'Accent', 'Dark2', 'Paired', 'Pastel1', 'Pastel2', 'Set1', 'Set2', 'Set3', 'tab10', 'tab20', 'tab20b', 'tab20c', 'magma\_r', 'inferno\_r', 'plasma\_r', 'viridis\_r', 'cividis\_r', 'twilight\_r', 'twilight\_shifted\_r', 'turbo\_r', 'Blues\_r', 'BrBG\_r', 'BuGn\_r', 'BuPu\_r', 'CMRmap\_r', 'GnBu\_r', 'Greens\_r', 'Greys\_r', 'OrRd\_r', 'Oranges\_r', 'PRGn\_r', 'PiYG\_r', 'PuBu\_r', 'PuBuGn\_r', 'PuOr\_r', 'PuRd\_r', 'Purples\_r', 'RdBu\_r', 'RdGy\_r', 'RdPu\_r', 'RdYlBu\_r', 'RdYlGn\_r', 'Reds\_r', 'Spectral\_r', 'Wistia\_r', 'YlGn\_r', 'YlGnBu\_r', 'YlOrBr\_r', 'YlOrRd\_r', 'afmhot\_r', 'autumn\_r', 'binary\_r', 'bone\_r', 'brg\_r', 'bwr\_r', 'cool\_r', 'coolwarm\_r', 'copper\_r', 'cubehelix\_r', 'flag\_r', 'gist\_earth\_r', 'gist\_gray\_r', 'gist\_heat\_r', 'gist\_ncar\_r', 'gist\_rainbow\_r', 'gist\_stern\_r', 'gist\_yarg\_r', 'gnuplot\_r', 'gnuplot2\_r', 'gray\_r', 'hot\_r', 'hsv\_r', 'jet\_r', 'nipy\_spectral\_r', 'ocean\_r', 'pink\_r', 'prism\_r', 'rainbow\_r', 'seismic\_r', 'spring\_r', 'summer\_r', 'terrain\_r', 'winter\_r', 'Accent\_r', 'Dark2\_r', 'Paired\_r', 'Pastel1\_r', 'Pastel2\_r', 'Set1\_r', 'Set2\_r', 'Set3\_r', 'tab10\_r', 'tab20\_r', 'tab20b\_r', 'tab20c\_r'\] The image below show the palette name and the colors: ![matplotlib-colors](https://datascientyst.com/content/images/2022/11/matplotlib-colors.webp) The code to generate this image can be found on this link: [How to view all colormaps available in matplotlib?](https://stackoverflow.com/a/68317686?ref=datascientyst.com) ## 8\. Conclusion We saw how to list huge amount of named colors in Python and Matplotlib. We saw also how to use named colors in Pandas. Finally we learned how to convert different color formats in Python and Pandas. Multiple examples of Matplotlib colors are shown. **This article can be used as a quick guide or reference for Matplotlib colors and Python styling.** ### How to Install Python and Pandas on Windows URL: https://datascientyst.com/install-python-pandas-windows/ Last updated: 2022-02-01T21:18:37.000Z ## 1\. Overview In this tutorial, we'll learn how to **install Pandas and Python on Windows**. We will cover the most popular ways of installation. The instructions described below have been tested on Windows 7 and 20. ## 2\. Install Python on Windows 10 Python is a widely-used easy to learn, user friendly, concise and high-level programming language. It is very easy to start coding on it and has a huge community. Python is one of the most liked and wanted languages according to: [stackoverflow - Python is the most wanted language for its fifth-year](https://insights.stackoverflow.com/survey/2021?ref=datascientyst.com#most-loved-dreaded-and-wanted-language-love-dread) ### 2.1\. Download and Install Python on Windows Unlike Linux, Windows doesn't come with a pre-installed Python version. So if we want to use the latest version of Python then manual installation is the way to go. The steps to install Python on Windows are: 1. Go to [Python Releases for Windows](https://www.python.org/downloads/windows/?ref=datascientyst.com) 2. Select the Python version you like - I prefer to go with the Stable Releases. For example - `Python 3.9.10 - Jan. 14, 2022` 3. Select the type of the installation - I prefer `Download Windows installer (64-bit)` \- [python-3.9.10-amd64.exe](https://www.python.org/ftp/python/3.9.10/python-3.9.10-amd64.exe?ref=datascientyst.com) 4. Download the desired version 5. Run the Python Installer 6. During installation 6.1\. Select **Install launcher for all users** \- if you like to have it for all users 6.2\. Check **Add Python 3.9 to PATH** in order Python to be visible for other programs 7. Press **Install Now** 8. Finally you can verify the installation by checking the version - `Python -V` \- expected output - `Python 3.9.10` ### 2.2\. Create a virtual environment (optional) Python offers a powerful package system `venv` which helps separate different Python packages. In simple words, you can create several virtual environments in multiple Pandas versions: - pandas 1.3.4 - pandas 1.0.0 To create new virtual environment called `pandas1`: - create folder for your virtual environments ( or select existing one) - Run command: `python -m venv pandas1` - activate the environment by: ```bash cd pandas1 source bin/activate ``` Once environment is activated you will see change in the terminal: ```python (pandas1) $ deactivate ``` The command above deactivates the environment. To learn more please check: - [venv — Creation of virtual environments](https://docs.python.org/3/library/venv.html?ref=datascientyst.com) ### 2.3\. Install PIP (optional) PIP is one of the most popular package managers for Python. If it's not installed: ```python pip -V ``` will return message that the command is not recognized - then you can install it by: ```python py -m ensurepip --upgrade ``` To learn more please check: - [PIP Installation](https://pip.pypa.io/en/stable/installation/?ref=datascientyst.com) ## 3\. Install Pandas on Windows ### 3.1\. Install Pandas by Pypi Next step is to install Pandas on Windows. The most easiest way of installing Pandas is by running: ```python pip install pandas ``` You can find more information for Pandas on: [pandas - pypi.org](https://pypi.org/project/pandas/?ref=datascientyst.com). ### 3.2\. Install Anaconda and Pandas on Windows If you like to use alternative installation methods you can check the official docs: [Installing Anaconda on Windows](https://docs.anaconda.com/anaconda/install/windows/?ref=datascientyst.com). For example Pandas is part of [Anaconda](https://docs.continuum.io/anaconda/?ref=datascientyst.com) \- so if you install Anaconda on your system you will get Pandas: - download the [Anaconda installer for Windows](https://www.anaconda.com/download/?ref=datascientyst.com#windows) - Verify data integrity with SHA-256\. (optional but highly RECOMMENDED step) - Install Anaconda by double click on the installer - Press Next - "I Agree" on - licensing terms - Select "Just Me" - if you are going to use it for yourself only - Select a destination folder - Click the Next button - Add "Anaconda to your PATH environment variable" - highly recommended - Continue the rest depending on your personal preferences or refer to the official installation guide ### 3.3\. Verify Pandas installation Finally you can test Pandas installation by running next commands: ```bash pip freeze | grep pandas ``` result will be: ```bash pandas==1.4.0 ``` ## 4\. Conclusion To summarize, in this article, we've seen examples of installing Python and Pandas on Windows in several ways. We've briefly explained these installation methods and how to verify the installation. And finally, we've seen how to manage multiple Python/Pandas installations on Windows like systems with different package versions. ### Pandas Most Typical Errors and Solutions for Beginners URL: https://datascientyst.com/pandas-most-typical-errors-and-solutions/ Last updated: 2023-03-17T23:05:23.000Z In this post I'll try to list the **most often errors and their solution in Pandas and Python.** The list will grow with time and will be updated frequently. ## DateTime ### Invalid comparison between or subtraction must have the same timezones - **`TypeError: Timestamp subtraction must have the same timezones or no timezones`** - **`datetimearray subtraction must have the same timezones or no timezones`** - **`TypeError: Invalid comparison between dtype=datetime64[ns] and DatetimeArray`** - **`TypeError: Invalid comparison between dtype=datetime64[ns] and Date`** Quick solution is to remove the timezone information by: ```python df['time_tz'].dt.tz_localize(None) ``` Example and more details: [How to Remove Timezone from a DateTime Column in Pandas](https://datascientyst.com/remove-timezone-datetime-column-pandas/) ### TypeError: Addition/subtraction of integers and integer-arrays with Timestamp is no longer supported - **`` TypeError: Addition/subtraction of integers and integer-arrays with Timestamp is no longer supported. Instead of adding/subtracting `n`, use `n * obj.freq` ``** Quick solution is to use `n * obj.freq`: ```python pd.to_datetime("today") - pd.Timedelta(10, unit='D') ``` Example and more details: [How to Get Today's Date in Pandas](https://datascientyst.com/get-todays-date-pandas/) ### 'index' object has no attribute 'tz\_localize' - **`'index' object has no attribute 'tz_localize'`** - **`attributeerror: 'index' object has no attribute 'tz_localize'`** Quick solution is to check if the index is from DateTime or convert a column before using it as index: ```python df.set_index(pd.DatetimeIndex(df['date']), drop=False, inplace=True) ``` Example and more details: [How to Remove Timezone from a DateTime Column in Pandas](https://datascientyst.com/remove-timezone-datetime-column-pandas/) ### OutOfBoundsDatetime - **`OutOfBoundsDatetime: Out of bounds nanosecond timestamp`** The short answer of this error is: ```python pd.to_datetime(df['date'], errors = 'ignore') ``` Example and more details: [OutOfBoundsDatetime: Out of bounds nanosecond timestamp - Pandas and pd.to\_datetime](https://datascientyst.com/outofboundsdatetime-out-of-bounds-nanosecond-timestamp-pandas-pd-to%5Fdatetime/) ### Wrong dates - ParserError: Unknown string format - **`ParserError: Unknown string format: 1975-02-23T02:58:41.000Z 1975-02-23T02:58:41.000Z`** ```python df['date'] = pd.to_datetime(df['date_str'], format='%d/%m/%Y', errors='coerce') ``` Example and more details: - [How to Fix Pandas to\_datetime: Wrong Date and Errors](https://datascientyst.com/how-to-fix-pandas-to%5Fdatetime-wrong-date-and-errors/) - [Combine Multiple columns into a single one in Pandas](https://datascientyst.com/combine-multiple-columns-into-single-one-in-pandas) ### Wrong dates - ValueError: time data does not match format '%Y%m%d HH:MM:SS' (match) - **`ValueError: time data '28-01-2022 5:25:00 PM' does not match format '%Y%m%d HH:MM:SS' (match)`** ```python pd.to_datetime('20220701', format='%Y%m%d', errors='ignore') ``` Example and more details: - [How to Convert String to DateTime in Pandas](https://datascientyst.com/convert-string-to-datetime-pandas/) ![dependencies](https://datascientyst.com/content/images/2022/03/dependencies.png) ## read\_csv ### UnicodeDecodeError - 'utf-8' codec can't decode byte 0x97 in position 6785: invalid start byte - **`UnicodeDecodeError: 'utf-8' codec can't decode byte 0x97 in position 6785: invalid start byte `** The short answer of this error is: ```python df = pd.read_csv('../data/csv/file_utf-16.csv', encoding='utf-16') ``` Example and more details: [How to Fix - UnicodeDecodeError: invalid start byte - during read\_csv in Pandas](https://datascientyst.com/pandas-read%5Fcsv-unicodedecodeerror-invalid-start-byte/) ### ParserError: Expected 5 fields in line 5, saw 6 - **`ParserError: Expected 5 fields in line 5, saw 6. Error could possibly be due to quotes being ignored when a multi-char delimiter is used`** The short answer of this error is: ```python df = pd.read_csv(csv_file, delimiter=';;', engine='python', error_bad_lines=False) df = pd.read_csv(csv_file, delimiter=';;', engine='python', on_bad_lines='skip) ``` Example and more details: [How to Use Multiple Char Separator in read\_csv in Pandas](https://datascientyst.com/use-multiple-char-separator-read%5Fcsv-pandas/) ### ParserError: Error tokenizing data. C error: Expected 2 fields in line 4, saw 4 - **`ParserError: Error tokenizing data. C error: Expected 2 fields in line 4, saw 4 `** The short answer of this error is: ```python pd.read_csv('test.csv', on_bad_lines='skip') ``` Example and more details: [How to Solve Error Tokenizing Data on read\_csv in Pandas](https://datascientyst.com/solve-error-tokenizing-data-read%5Fcsv-pandas/) ## to\_csv ### AttributeError: object has no attribute 'to\_csv' - **`AttributeError: 'numpy.ndarray' object has no attribute 'to_csv'`** - **`attributeerror: 'tuple' object has no attribute 'to_csv'`** The short answer of this error is: ```python pd.Series(df['Magnitude Type'].unique()).to_csv('data.csv') ``` Example and more details: [Dump (unique) values to CSV / to\_csv in Pandas](https://datascientyst.com/dump-unique-values-to-csv-to%5Fcsv-pandas/) ## Index / MultiIndex ### MultiIndex - indexerror: too many levels - **`ValueError: Cannot remove 1 levels from an index with 1 levels: at least one level must be left.`** - **`IndexError: Too many levels: Index has only 1 level, not 4`** - **`indexerror: too many levels: index has only 1 level, not 2`** The short answer of this error is: ```python df.index df.droplevel(level=1) df.reset_index(level=1) df.columns.droplevel(level=0) ``` Example and more details: [How to Drop a Level from a MultiIndex in Pandas DataFrame](https://datascientyst.com/pandas-drop-multiindex-level/) ### Sort MultiIndex - label must be a tuple with elements corresponding to each level - **`ValueError: The column label 'Depth' is not unique. For a multi-index, the label must be a tuple with elements corresponding to each level.`** The short answer of this error is: ```python df_multi.columns df_multi.columns.get_level_values(1) df_multi.sort_values(by=[('Depth', 'mean')], ascending=False) ``` Example and more details: [How to Sort MultiIndex in Pandas](https://datascientyst.com/sort-multiindex-pandas/) ## Merge ### ValueError: Indexes have overlapping values - **`ValueError: Indexes have overlapping values: Index(['A', 'B', 'C', 'D'], dtype='object')`** The short answer of this error is: ```python pd.concat([df1, df2], axis='columns', verify_integrity=False) df1.join(df2, lsuffix='_x') ``` Example and more details: [How to Merge Two DataFrames on Index in Pandas](https://datascientyst.com/merge-two-dataframes-on-index-pandas/) ## String ### TypeError: can only concatenate str (not "float") to str - **`TypeError: can only concatenate str (not "float") to str`** The short answer of this error is: ```python df['Magnitude'].astype(str) ``` Example and more details: [Combine Multiple columns into a single one in Pandas](https://datascientyst.com/combine-multiple-columns-into-single-one-in-pandas) ## Column ### ValueError: cannot reindex from a duplicate axis - **`ValueError: cannot reindex from a duplicate axis`** The short answer of this error is: ```python df = df.sort_index(axis=1) ``` Example and more details: [How to Change the Order of Columns in Pandas DataFrame](https://datascientyst.com/change-order-columns-pandas-dataframe) ## DataFrame ### ValueError: `caption` must be either a string or 2-tuple of strings. - **`ValueError: `caption` must be either a string or 2-tuple of strings.`** The short answer of this error is - Use string or 2-tuple for DataFrame captions - and not integer: ```python df.style.set_caption('DataFrame Name') ``` Example and more details: [How to Set Caption and Customize Font Size and Color in Pandas DataFrame](https://datascientyst.com/set-caption-customize-font-size-color-in-pandas-dataframe/) ### How to Set Caption and Customize Font Size and Color in Pandas DataFrame URL: https://datascientyst.com/set-caption-customize-font-size-color-in-pandas-dataframe/ Last updated: 2022-02-01T20:37:19.000Z ## 1\. Overview In this short guide we will see how to **set and customize the caption of the DataFrame styler in Pandas.** We are going to set a new caption, change the format: the font, the font size, the color etc. ## 2\. Setup For this guide we are going to create a simple DataFrame with the following data: ```python import pandas as pd data = {'Member': {0: 'John', 1: 'Bill'}, 'Disqualified': {0: 0, 1: 1}, 'Paid': {0: 1, 1: 0}} df = pd.DataFrame(data) ``` data looks like this: | | Member | Disqualified | Paid | | - | ------ | ------------ | ---- | | 0 | John | 0 | 1 | | 1 | Bill | 1 | 0 | ## 3\. Set new caption for DataFrame In order set new name for the DataFrame in Pandas we can use the following method [pandas.io.formats.style.Styler.set\_caption](https://pandas.pydata.org/docs/reference/api/pandas.io.formats.style.Styler.set%5Fcaption.html?ref=datascientyst.com). Setting new title can accept only string or 2-tuple: ```python df.style.set_caption('Members') ``` result after the DataFrame renaming: __Members__ | | Member | Disqualified | Paid | | - | ------ | ------------ | ---- | | 0 | John | 0 | 1 | | 1 | Bill | 1 | 0 | As you can see the If you try to use integer then error is raised: > ValueError: `caption` must be either a string or 2-tuple of strings. ## 4\. Customize the color, font size for caption for DataFrame To customize the color, font size and text alignment of the caption we can use the `set_table_styles()` method. Set: - set new color - lime - specify the font-size - 150% - set text-align - left ```python styles = [dict(selector="caption", props=[("text-align", "right"), ("font-size", "150%"), ("color", 'lime')])] df.style.set_caption('Members').set_table_styles(styles) ``` result: __Members__ | | Member | Disqualified | Paid | | - | ------ | ------------ | ---- | | 0 | John | 0 | 1 | | 1 | Bill | 1 | 0 | We are free to use HTML and CSS properties like: - `background-color` \- background color - `color` \- text colors - `font-size` \- text sizes - `text-align` \- text alignment - `font-family` \- text fonts To find what properties can be set you can read those two links: - [HTML Styles](https://www.w3schools.com/html/html%5Fstyles.asp?ref=datascientyst.com) - [HTML Text Formatting](https://www.w3schools.com/html/html%5Fformatting.asp?ref=datascientyst.com) ## 5\. Set bold or multiline title for DataFrame In this section we will cover more examples on customizing the DataFrame caption. ### 5.1\. Set bold caption for DataFrame To change the title or other properties of your DataFrame you can use: `("font-weight", "bold")`: ```python styles = [dict(selector="caption", props=[("font-size", "120%"), ("font-weight", "bold")])] df.style.set_caption('Members').set_table_styles(styles) ``` result: __Members__ | | Member | Disqualified | Paid | | - | ------ | ------------ | ---- | | 0 | John | 0 | 1 | | 1 | Bill | 1 | 0 | ### 5.2\. Set multiline caption for DataFrame We can also add HTML tags in the title like the new line HTML tag - `
`: ```python styles = [dict(selector="caption", props=[("font-size", "120%")])] df.style.set_caption('Members
52022').set_table_styles(styles) ``` result: __Members20202__ | | Member | Disqualified | Paid | | - | ------ | ------------ | ---- | | 0 | John | 0 | 1 | | 1 | Bill | 1 | 0 | ### 5.3\. Set background for DataFrame title - set\_caption() To change the back-ground color we can use the method `.set_table_styles(styles)`: ```python styles = [dict(selector="caption", props=[("background-color", "cyan")])] df.style.set_caption('Members').set_table_styles(styles) ``` The new DataFrame have different background: __Members__ | | topic\_1 | topic\_2 | topic\_3 | topic\_4 | | ------- | -------- | -------- | -------- | -------- | | item\_1 | a | nan | e | nan | | item\_2 | nan | c | nan | nan | | item\_3 | nan | d | nan | f | | item\_4 | nan | nan | nan | nan | | item\_5 | b | nan | nan | d | ## 6\. Get the HTML for the styled DataFrame Finally if you like to get the HTML code for your customized DataFrame we can use `print()` and `to_html()` methods: ```python print(df.style.set_caption('Members').set_table_styles(styles).to_html()) ``` ## 7\. Conclusion We saw how to set custom titles for DataFrame. Also we covered several examples about changing the font size, font, color and other attributes. ### How to Get First Non-NaN Value Per Row in Pandas URL: https://datascientyst.com/get-first-non-null-value-per-row-pandas/ Last updated: 2022-06-21T06:41:18.000Z ## 1\. Overview To get the **first or the last non-NaN value per rows in Pandas** we can use the next solutions: **(1) Get First/Last Non-NaN Values per row** ```python df.fillna(method='bfill', axis=1).iloc[:, 0] ``` **(2) Get non-NaN values with stack() and groupby()** ```python df.stack().groupby(level=0).first().reindex(df.index) ``` **(3) Get the column name of first non-NaN value per row** - If you like to learn more about this problem please check: [Pandas Combine Multiple Columns Into a Single in Pandas](https://datascientyst.com/combine-multiple-columns-into-single-one-in-pandas/) ```python df.apply(pd.Series.first_valid_index, axis=1) ``` In the next steps we will cover all the examples in detail. ## 2\. Setup ### 2.1\. Data For this tutorial we are going to use the following DataFrame: ```python import pandas as pd import numpy as np details = { 'topic_1': {'item_1': 6, 'item_2': np.NaN, 'item_3': np.NaN, 'item_4': np.NaN, 'item_5': 5}, 'topic_2': {'item_1': np.NaN, 'item_2': 8, 'item_3': 5, 'item_4': np.NaN, 'item_5': np.NaN}, 'topic_3': {'item_1': 2, 'item_2': np.NaN, 'item_3': np.NaN, 'item_4': np.NaN, 'item_5': np.NaN}, 'topic_4': {'item_1': np.NaN, 'item_2': np.NaN, 'item_3': 7, 'item_4': np.NaN, 'item_5': 6} } df = pd.DataFrame(details) ``` Data looks like: | | topic\_1 | topic\_2 | topic\_3 | topic\_4 | | ------- | -------- | -------- | -------- | -------- | | item\_1 | 6.0 | NaN | 2.0 | NaN | | item\_2 | NaN | 8.0 | NaN | NaN | | item\_3 | NaN | 5.0 | NaN | 7.0 | | item\_4 | NaN | NaN | NaN | NaN | | item\_5 | 5.0 | NaN | NaN | 6.0 | ### 2.2\. Expected Result The expectation is to get first non-NaN values per given row: ``` item_1 6.0 item_2 8.0 item_3 5.0 item_4 NaN item_5 5.0 Name: topic_1, dtype: float64 ``` ## 3\. Get First/Last Non-NaN Values per row The first solution to **get the non-NaN values per row from a list of columns** use the next steps: - `.fillna(method='bfill', axis=1)` \- to fill all non-NaN values from the last to the first one; `axis=1` \- means columns - `.iloc[:, 0]` \- get the first column So the final code will looks like: ```python df.fillna(method='bfill', axis=1).iloc[:, 0] ``` and the result Series will have all non-null values per given row: ``` item_1 6.0 item_2 8.0 item_3 5.0 item_4 NaN item_5 5.0 Name: topic_1, dtype: float64 ``` To **get the last non-NaN value per row** you need to change the code to: ```python df.fillna(method='ffill', axis=1).iloc[:, -1] ``` and the result would be: ``` item_1 2.0 item_2 8.0 item_3 7.0 item_4 NaN item_5 6.0 Name: topic_4, dtype: float64 ``` Note: To work only with the needed columns we can select them as `df[['topic_1', 'topic_2']]` ## 4\. Get non-NaN values with stack() and groupby() Another option to get the first or last non-NaN values is by combination of methods `stack()` and `groupby()`. The idea is to dynamically create a single column with first non-NaN value: ```python df.stack().groupby(level=0).first().reindex(df.index) ``` result: ``` item_1 6.0 item_2 8.0 item_3 5.0 item_4 NaN item_5 5.0 dtype: float64 ``` The following algorithm explains how the code works: - Use method `stack()` in order to stack the values: item\_1 topic\_1 6.0 topic\_3 2.0 item\_2 topic\_2 8.0 item\_3 topic\_2 5.0 topic\_4 7.0 item\_5 topic\_1 5.0 topic\_4 6.0 dtype: float64 - `.groupby(level=0).first()` \- will group by the level 0 of the multi-index and return the first values. The result would be: item\_1 6.0 item\_2 8.0 item\_3 5.0 item\_5 5.0 dtype: float64 - `.reindex(df.index)` \- the final step is to reindex values on the original DataFrame in order to add all missing values: item\_1 6.0 item\_2 8.0 item\_3 5.0 item\_4 NaN item\_5 5.0 dtype: float64 ## 5\. Get the column name of first non-NaN value per row If you need the column names instead of the values we can use the following code: ```python df.apply(pd.Series.first_valid_index, axis=1) ``` which will **return the column name per each row which has the first non-NaN values:** ``` item_1 topic_1 item_2 topic_2 item_3 topic_2 item_4 None item_5 topic_1 dtype: object ``` The method `first_valid_index()` works as follows: > Return index for first non-NA value or None, if no NA value is found. ## 6\. Conclusion We covered how to work with multiple columns and NaN values. Now we know how to get the first or last non-empty values per row. We also described and combined several Pandas functions like: - [pandas.DataFrame.stack](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.stack.html?ref=datascientyst.com) - [pandas.DataFrame.fillna](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.fillna.html?ref=datascientyst.com) - [pandas.DataFrame.groupby](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.groupby.html?ref=datascientyst.com) - [pandas.DataFrame.first](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.first.html?ref=datascientyst.com) - [pandas.DataFrame.reindex](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.reindex.html?ref=datascientyst.com) - [pandas.DataFrame.first\_valid\_index](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.first%5Fvalid%5Findex.html?ref=datascientyst.com) Finally we show how to understand the logic behind complex Pandas code. ### How to Get Column Name of First Non NaN value in Pandas URL: https://datascientyst.com/get-column-name-first-non-nan-value-pandas/ Last updated: 2022-04-13T08:33:08.000Z In this tutorial, we'll learn **how to select non null values in Pandas**. You can easily find also the number of non NaN or NaN values in column or multiple columns. In the next section we will cover all the steps in a real world example. To learn more about the NaN values in Pandas you can check also: [How to Get First Non-NaN Value Per Row in Pandas](https://datascientyst.com/get-first-non-null-value-per-row-pandas/). ## Setup Let's start with sample data and the problem for this tutorial. ### DataFrame with NaN values We have different items on each row. The columns represent different topics. One item can have zero or multiple topics Our goal is to get the first or last topic for each item. ```python import pandas as pd details = { 'topic_1': {'item_1': 1, 'item_2': 0, 'item_3': 0, 'item_4': 0, 'item_5': 1}, 'topic_2': {'item_1': 0, 'item_2': 1, 'item_3': 1, 'item_4': 0, 'item_5': 0}, 'topic_3': {'item_1': 1, 'item_2': 0, 'item_3': 0, 'item_4': 0, 'item_5': 0}, 'topic_4': {'item_1': 0, 'item_2': 0, 'item_3': 1, 'item_4': 0, 'item_5': 0} } df = pd.DataFrame(details) ``` result: | | topic\_1 | topic\_2 | topic\_3 | topic\_4 | | ------- | -------- | -------- | -------- | -------- | | item\_1 | 1 | 0 | 1 | 0 | | item\_2 | 0 | 1 | 0 | 0 | | item\_3 | 0 | 1 | 0 | 1 | | item\_4 | 0 | 0 | 0 | 0 | | item\_5 | 1 | 0 | 0 | 0 | ### Intro - Get Name of First Non NaN Column Can you extract a topic for each row? The problem is to identify the first or last topic from multiple for each item. Expected result: | | topic\_1 | topic\_2 | topic\_3 | topic\_4 | category | | ------- | -------- | -------- | -------- | -------- | -------- | | item\_1 | 1 | 0 | 1 | 0 | topic\_3 | | item\_2 | 0 | 1 | 0 | 0 | topic\_2 | | item\_3 | 0 | 1 | 0 | 1 | topic\_4 | | item\_4 | 0 | 0 | 0 | 0 | other | | item\_5 | 1 | 0 | 0 | 0 | topic\_1 | ## Step 1: Get Column name for First non NaN Column Let's start by **getting the first topic per row/item from a list of columns**. First we will identify a list of columns which are going to be used. We have two options: - get the columns by `df.columns` - or write them explicitly - `categories = reversed(['topic_1', 'topic_2', 'topic_3', 'topic_4'])` We are going to use simple algorithm to get values from the multiple columns: - create new column with default value of 'other' - iterate over all rows - iterate over all categories - map the values to the topic name - fill the missing values from the last values of the new column - this step is needed in order to replace non matched values which will be assigned with NaNs. The code below is **getting the first non NaN values from each column**: ```python categories = reversed(['topic_1', 'topic_2', 'topic_3', 'topic_4']) df['category'] = 'other' for ix, row in df.iterrows(): for cat in categories: d = {1: cat} df['category'] = df[cat].map(d).fillna(df['category']) ``` To get the last non NaN value we can change the list order by removing `reversed`. ## Step 2: Get columns with NaN values Pandas - explanation In this step we will describe how the main part of the code is working. The main part has two important functions: - [pandas.Series.map](https://pandas.pydata.org/docs/reference/api/pandas.Series.map.html?ref=datascientyst.com) \- maps a dict to a column and returns all found values. Otherwise returns NaN. You can learn more in this article: [How to Map Column with Dictionary in Pandas](https://datascientyst.com/pandas-map-column-dictionary/) - [pandas.DataFrame.fillna](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.fillna.html?ref=datascientyst.com) \- fills all missing values with the provided value or iterable(in this case) So the idea is to map all values from the current column and with next columns to fill the new values while keeping the ones which are set. ## Conclusion In this short article we saw how to combine multiple Pandas functions in order to achieve complex logic - **get first/last column from multiple which has non NaN value**. Using the described technique you can combine multiple columns to a single one based on their values. ### How to update Pandas in PIP, Anaconda, Poetry URL: https://datascientyst.com/update-pandas-pip-anaconda-poetry/ Last updated: 2022-01-30T10:18:42.000Z To **update Pandas to a specific or to a latest version** you need to check next steps: ## Steps to update Pandas - Activate the environment for Pandas - Check how Pandas was installed - Pip - Anaconda - Poetry - Check current Pandas version - optional - Upgrade Pandas to - specific version - latest version ## Step 2\. Check how Pandas was installed To start you need to know how Pandas is installed and where. The simplest but not best way is to check which command is working for you: **To list packages in Pip virtual environment** ```python pip list ``` or to **check for Anaconda environment** \- get a list of all conda environments ```python conda info --envs ``` To check Poetry environment use: ```python poetry show --latest ``` There is a good extra reading on the topic here with best practices: [Using Pip in a Conda Environment](https://www.anaconda.com/blog/using-pip-in-a-conda-environment?ref=datascientyst.com) ## Step 3\. Check current Pandas version To check which is the current Pandas version you can use: ### 3.1\. Check Pandas version by PIP ```python !pip list | grep pandas ``` which will show something like: ``` pandas 1.4.0 ``` or the more detailed information by: ```python !pip show pandas ``` ``` Name: pandas Version: 1.4.0 Summary: Powerful data structures for data analysis, time series, and statistics Home-page: https://pandas.pydata.org Author: The Pandas Development Team Author-email: pandas-dev@python.org License: BSD-3-Clause Location: /home/user/Software/Tensorflow/environments/venv38/lib/python3.8/site-packages Requires: numpy, python-dateutil, pytz Required-by: altair, bar-chart-race, calmap, calplot, dataframe-image, geopandas, holoviews, ibis-framework, itables, modin, pyLDAvis, qgrid, seaborn, statsmodels, swifter, tabula-py Note: you may need to restart the kernel to use updated packages ``` ### 3.2\. Check current Pandas version in Conda To check which is the installed Pandas version in Conda use the following: ```python conda list pandas ``` To find all available versions use: ```python conda search pandas ``` ### 3.3\. Check Pandas package in poetry ```python poetry show pandas ``` result: name : pandas version : 1.4.0 ## 4\. Update Pandas ### 4.1\. Update Pandas with Pip To install specific version of Pandas by Pip use this format: ```python pip install pandas==1.3.2 ``` and for upgrade to the latest version use: ```python pip install --upgrade pandas ``` ### 4.2\. Update Pandas in Conda To install specific version of Pandas with Anaconda use this format: ```python conda install pandas=1.0.2 ``` and for update to the latest version use: ```python conda install pandas ``` ### 4.3\. Update Pandas in Poetry To update Pandas to the latest version in Poetry you can use: ```python poetry update pandas ``` for a specific version use: ```python poetry add pandas="1.3.2" ``` Depending on the `.toml` file different versions can be installed. To read more about poetry follow this link: [Poetry update](https://python-poetry.org/docs/cli/?ref=datascientyst.com#update) ### How to Create a DataFrame from Lists in Pandas URL: https://datascientyst.com/create-dataframe-from-lists-pandas/ Last updated: 2022-01-30T09:23:29.000Z To create a DataFrame from list or nested lists in Pandas we can follow next steps: ## Steps - Define the input data - flat list - nested list - dict of lists - Define the column names - Define the index - row names - Call the DataFrame constructor ## Examples In several examples we will cover the most popular cases of using lists to create DataFrames ### Example 1 - Create Pandas DataFrame from List Starting with a flat list which will be used for the values of the Future DataFrame. We are providing the index and the columns - which are optional. If the index and columns are not specified - then automatic ones will be used. ```python import pandas as pd values = ['val_1', 'val_2', 'val_3'] df = pd.DataFrame(values, index = ['x', 'y', 'z'], columns = ['col_1']) ``` ### Output 1 | | col\_1 | | - | ------ | | x | val\_1 | | y | val\_2 | | z | val\_3 | ### Example 2 - Create Pandas DataFrame from Nested Lists The second example demonstrates creating a multi column DataFrame from nested lists. We can specify the column names as follows: ```python import pandas as pd values = [['val_11', 'val_12'], ['val_21', 'val_22'], ['val_31', 'val_32']] df = pd.DataFrame(values, columns = ['col_1', 'col_2']) ``` ### Output 2 | | col\_1 | col\_2 | | - | ------- | ------- | | 0 | val\_11 | val\_12 | | 1 | val\_21 | val\_22 | | 2 | val\_31 | val\_32 | ### Example 3 - Create Pandas DataFrame from dict of Lists In this example we are creating multiple columns from several lists. Those lists are mapped with a dictionary to different columns. If you want to learn more about creating DataFrames from dictionaries you can read this detailed article - [How to Create DataFrame from Dictionary in Pandas?](https://datascientyst.com/create-dataframe-from-dictionary-pandas/) ```python import pandas as pd col_1 = ["val_11", "val_21", "val_31"] col_2 = ["val_21", "val_22", "val_32"] data = {'col_1': col_1, 'col_2': col_2} df = pd.DataFrame(data) ``` ### Output 3 | | col\_1 | col\_2 | | - | ------- | ------- | | 0 | val\_11 | val\_21 | | 1 | val\_21 | val\_22 | | 2 | val\_31 | val\_32 | ### How to Create DataFrame from Dictionary in Pandas? URL: https://datascientyst.com/create-dataframe-from-dictionary-pandas/ Last updated: 2022-01-30T08:58:54.000Z ## 1\. Overview To **create DataFrame from dictionary in Pandas** there are several options depending on the dictionary and desired result: **(1) create DataFrame from dictionary using default Constructor** ```python data = {'Member': {0: 'John', 1: 'Bill', 2: 'Jim'}, 'Disqualified': {0: 0, 1: 1, 2: 0}, 'Paid': {0: 1, 1: 0, 2: 3}} df = pd.DataFrame(data) ``` result: | | Member | Disqualified | Paid | | - | ------ | ------------ | ---- | | 0 | John | 0 | 1.0 | | 1 | Bill | 1 | 0.0 | | 2 | Jim | 0 | 3.0 | **(2) use method from\_dict() to create DataFrame from dictionary** ```python import pandas as pd data = {'col_1': [1, 2, 3], 'col_2': ['x', 'y', 'z']} pd.DataFrame.from_dict(data) ``` result: | | col\_1 | col\_2 | | - | ------ | ------ | | 0 | 1 | x | | 1 | 2 | y | | 2 | 3 | z | Let's cover those ways in more detail in the next steps. ## 2\. Use method from\_dict() to create DataFrame from dictionary Explicitly using the method [from\_dict()](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.from%5Fdict.html?ref=datascientyst.com) to create DataFrame from dictionary has several options: - default behavior - the keys of the dict will be the columns - `orient='index'` \- dictionary keys as rows - `orient='tight'` \- using a tight format ### 2.1\. Pandas from\_dict() - default behavior By default the method `from_dict()` will use the keys of the dictionary as column names. The dictionary values will be used as a DataFrame values: ```python import pandas as pd data = {'col_1': [1, 2, 3], 'col_2': ['x', 'y', 'z']} df = pd.DataFrame.from_dict(data) ``` result: | | col\_1 | col\_2 | | - | ------ | ------ | | 0 | 1 | x | | 1 | 2 | y | | 2 | 3 | z | ### 2.2\. Pandas from\_dict() - dictionary keys as rows If you like to use dict keys as a rows you can specify parameter `orient='index'` \- this will create DataFrame with values from the dict values: ```python import pandas as pd data = {'row_1': [1, 2, 3], 'row_2': ['x', 'y', 'z']} df = pd.DataFrame.from_dict(data, orient='index') ``` the new DataFrame will looks like: | | 0 | 1 | 2 | | ------ | - | - | - | | row\_1 | 1 | 2 | 3 | | row\_2 | x | y | z | ### 2.3\. Pandas from\_dict() - using a tight format The final option is to use tight format - `orient='tight'`. This option is available only with version 1.4 - otherwise error will be raised: > ValueError: only recognize index or columns for orient It can used to create **MultiIndex DataFrames from dictionary like:** ```python data = {'index': [('row_1', 'row_21'), ('row_1', 'row_22')], 'columns': [('col_11', 'col_21'), ('col_12', 'col_22')], 'data': [['val_11', 'val_21'], ['val_12', 'val_22']], 'index_names': ['ix level 1', 'ix level 2'], 'column_names': ['col level 1', 'col level 2']} pd.DataFrame.from_dict(data, orient='tight') ``` result: | | col level 1 | col\_11 | col\_12 | | ---------- | ----------- | ------- | ------- | | | col level 2 | col\_21 | col\_22 | | ix level 1 | ix level 2 | | | | row\_1 | row\_21 | val\_11 | val\_21 | | row\_22 | val\_12 | val\_22 | | ## 3\. Create DataFrame from dictionary using default Constructor One more way to create DataFrame from dicts is by using the Constructor and providing the dictionary to it. There are several different options: - **Dict with list values** \- dict keys as columns and dict values as data - default index - explicit index - **Nested dictionary to create DataFrame** \- define index for the rows ### 3.1\. Dict with list values If the input dict has list values than they will be used as values for the DataFrame: ```python data = {'Member': ['John', 'Bill', 'Jim'], 'Disqualified': [0,1,2]} df = pd.DataFrame(data) ``` the output is: | | Member | Disqualified | | - | ------ | ------------ | | 0 | John | 0 | | 1 | Bill | 1 | | 2 | Jim | 2 | To define index in this case we can pass it to the parameter by: ```python pd.DataFrame(data, index = ['x', 'y', 'z']) ``` ### 3.2\. Nested dictionary to create DataFrame When we are using nested dictionary we will use index:value pairs for the dict: ```python details = { 'col_1' : { 'row_1' : 'x', 'row_2' : 'y', 'row_3' : 'z' }, 'col_2' : { 'row_1' : 1, 'row_2' : 2, 'row_3' : 3 } } df = pd.DataFrame(details) ``` result: | | col\_1 | col\_2 | | ------ | ------ | ------ | | row\_1 | x | 1 | | row\_2 | y | 2 | | row\_3 | z | 3 | ## 4\. Conclusion To summarize, in this article,\*\* we've seen examples of creating DataFrame from dictionary in multiple ways\*\*. We've briefly discussed these creation methods. And finally, we've seen how to create a MultiIndex DataFrame from a dict. ### 325-map URL: https://datascientyst.com/pandas-map/ Last updated: 2022-02-27T08:33:02.000Z map() ### How to fix SettingWithCopyWarning in Pandas: A value is trying to be set on a copy URL: https://datascientyst.com/fix-settingwithcopywarning-pandas-value-trying-set-copy/ Last updated: 2023-04-10T08:06:10.000Z ## 1\. Overview In this tutorial, we'll learn **how to solve the popular warning in Pandas and Python - SettingWithCopyWarning**: > /tmp/ipykernel\_4904/714243365.py:1: SettingWithCopyWarning: > A value is trying to be set on a copy of a slice from a DataFrame. > Try using .loc\[row\_indexer,col\_indexer\] = value instead > See the caveats in the documentation: [https://pandas.pydata.org/pandas-docs/stable/user\_guide/indexing.html#returning-a-view-versus-a-copy](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/indexing.html?ref=datascientyst.com#returning-a-view-versus-a-copy) Several different reasons can cause this warning message. We are going to cover most of them and their solutions. × If you face warnings in Pandas try to understand and resolve them. Ignoring or skipping might result in unexpected behaviour. ## 2\. Setup For this example we are going to use dummy [DataFrame created by method](https://datascientyst.com/create-easily-dummy-dataframe-test-data/): `makeMixedDataFrame`: ```python from pandas.util.testing import makeMixedDataFrame df = makeMixedDataFrame() ``` data: | | A | B | C | D | | - | --- | --- | ---- | ---------- | | 0 | 0.0 | 0.0 | foo1 | 2009-01-01 | | 1 | 1.0 | 1.0 | foo2 | 2009-01-02 | | 2 | 2.0 | 0.0 | foo3 | 2009-01-05 | | 3 | 3.0 | 1.0 | foo4 | 2009-01-06 | | 4 | 4.0 | 0.0 | foo5 | 2009-01-07 | ## 3\. What are the reasons for SettingWithCopyWarning ### 3.1\. Is Pandas DataFrame a Copy or a View? Before jumping to solutions, let's try to answer the question in the title of this section. How to tell the difference between Copy or a View? Let's cover this in few examples showing how to copy a DataFrame or get some part of it: ```python df_2 = df df_3 = df.copy() df_4 = df[:] df_5 = df.loc[:, :] df_6 = df.iloc[0:2, :] df_7 = df['D'] ``` Let's verify which of them are copies and which are views: ```python print(df._is_view, '|', hex(id(df)), '|', df._is_copy) print(df_2._is_view, '|', hex(id(df_2)), '|', df_2._is_copy) print(df_3._is_view, '|', hex(id(df_3)), '|', df_3._is_copy) print(df_4._is_view, '|', hex(id(df_4)), '|', df_4._is_copy) print(df_5._is_view, '|', hex(id(df_5)), '|', df_5._is_copy) print(df_6._is_view, '|', hex(id(df_6)), '|', df_6._is_copy) print(df_7._is_view, '|', hex(id(df_7)), '|', df_7._is_copy) ``` The output helps us to understand Copies and Views better: | \_is\_view | hex(id( | \_is\_copy | | ---------- | -------------- | ------------------------------------------------------------- | | False | 0x7fde7ab88eb0 | None | | False | 0x7fde7ab88eb0 | None | | False | 0x7fde7abf24f0 | None | | False | 0x7fdec136f940 | | | False | 0x7fde7ab88eb0 | None | | False | 0x7fde7abf2730 | | | True | 0x7fde7abf2340 | None | So we can see that: `df['D']` will return a copy. `df.iloc[0:2, :]` and `df[:]` returns views. We can also see the addresses of all DataFrames. One more way to check the values of the DataFrame is by attribute: ```python df_2.values.base ``` which will show difference in case of different values: ``` array([[0.0, 1.0, 2.0, 3.0, 4.0], [0.0, 1.0, 0.0, 1.0, 0.0], ['foo1', 'foo2', 'foo3', 'foo4', 'foo5'], [Timestamp('2009-01-01 00:00:00'), Timestamp('2009-01-02 00:00:00'), Timestamp('2009-01-05 00:00:00'), Timestamp('2009-01-06 00:00:00'), Timestamp('2009-01-07 00:00:00')]], dtype=object) ``` In some cases using `df_2.values` might lead to controversial results. ### 3.2\. How data is accessed Depending on how DataFrame data is accessed - will result in showing a warning or not. Let's say that we would like to update values in column `C`. We can do this by: ```python df["C"][df["C"]=="foo3"] = "foo33" ``` but warning will be produced: ``` /tmp/ipykernel_9907/1845991504.py:1: SettingWithCopyWarning: A value is trying to be set on a copy of a slice from a DataFrame See the caveats in the documentation: https://pandas.pydata.org/pandas-docs/stable/user_guide/indexing.html#returning-a-view-versus-a-copy df["C"][df["C"]=="foo3"] = "foo33" ``` To get and set the values without `SettingWithCopyWarning` warning we need to use `loc`: ```python df.loc[df["C"]=="foo3", "C"] = "foo333" ``` ## 4\. Fix SettingWithCopyWarning by method copy() The first and simplest solution is to create a DataFrame copy and work with it. This can be done by method - `copy()`. ### 4.1 a value is trying to be set on a copy of a slice from a dataframe Let's do a short demo of this problem: ``` /tmp/ipykernel_4904/714243365.py:1: SettingWithCopyWarning: A value is trying to be set on a copy of a slice from a DataFrame. Try using .loc[row_indexer,col_indexer] = value instead ``` and the solution. Let say that we get part of the initial DataFrame by: ```python df_new = df[['D', 'B']] ``` Our goal is to work only with this subset of columns and create new column based on the existing ones: ```python df_new['E'] = df_new['B'] > 0 ``` This will cause warning: ``` df_7['E'] = df_7['B'] > 0 df_7['E'] = df_7['B'] > 0 /tmp/ipykernel_9907/381168311.py:1: SettingWithCopyWarning: A value is trying to be set on a copy of a slice from a DataFrame. Try using .loc[row_indexer,col_indexer] = value instead See the caveats in the documentation: https://pandas.pydata.org/pandas-docs/stable/user_guide/indexing.html#returning-a-view-versus-a-copy df_7['E'] = df_7['B'] > 0 ``` This warning can be suppress by using method `copy()`: ```python df_new = df[['D', 'B']].copy() ``` **Note:** Using method - \`copy()\` is recommended to small and medium sized DataFrames. For big ones and production solutions will cause performance issues. ## 5\. Fix SettingWithCopyWarning by method loc In this section we will do a demo on the warning when we work with a single DataFrame. In this case the warning is caused by the way we access data. For example if we like to change all values in column `C` which are different from `foo3` then we might use: ```python df["C"][df["C"]!="foo3"] = "foo" ``` This will raise the warning message: > SettingWithCopyWarning: > A value is trying to be set on a copy of a slice from a DataFrame To perform the operation without `SettingWithCopyWarning` \- we need to use attribute `loc` in this way: ```python df.loc[df["C"]!="foo3", "C"] = "foo" ``` **the operation is completed without the warning: "try using .loc\[row\_indexer,col\_indexer\] = value instead"** ## 6\. When .loc results into SettingWithCopyWarning In some cases even if we are using recommendation of `df.loc[:, 'm']` we might get the error like: ```python df.loc[:, 'm'] = df['date'].dt.to_period('M') ``` raise an error: > See the caveats in the documentation: [https://pandas.pydata.org/pandas-docs/stable/user\_guide/indexing.html#returning-a-view-versus-a-copy](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/indexing.html?ref=datascientyst.com#returning-a-view-versus-a-copy) > df.loc\[:, 'm'\] = df\['date'\].dt.to\_period('M') > /tmp/ipykernel\_12403/2549023945.py:3: SettingWithCopyWarning: > A value is trying to be set on a copy of a slice from a DataFrame. > Try using .loc\[row\_indexer,col\_indexer\] = value instead To solve the error we need explicitly to add `copy()` method: ```python df.loc[:, 'm'] = df['date'].copy().dt.to_period('M') ``` Finally error **SettingWithCopyWarning** is solved. ## 7\. suppressing SettingWithCopyWarning warning Sometimes you may get the **error for SettingWithCopyWarning** when the code is correct and valid. Second execution of the problematic cell might hide the wanrning. To **get rid completely of this warning SettingWithCopyWarning** we can use method: `filterwarnings('ignore')` ```python import numpy as np np.warnings.filterwarnings('ignore') ``` ## 8\. Conclusion In this article, we looked at the **reasons and solutions for the warning SettingWithCopyWarning in Pandas.** We focused on solving the original cause of the: > "a value is trying to be set on a copy of a slice from a dataframe. try using .loc\[row\_indexer,col\_indexer\] = value instead". We also covered how to suppress the message itself. Update of Pandas library to the latest version might help to reduce the error. Finally depending on your context you may get a slightly different problem. For example working with multi-index which is explained here: [Returning a view versus a copy](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/indexing.html?ref=datascientyst.com#returning-a-view-versus-a-copy) ![](https://datascientyst.com/content/images/2022/01/fix-settingwithcopywarning-pandas-value-trying-set-copy.png) ### How to select rows by column value in Pandas URL: https://datascientyst.com/select-rows-column-value-pandas/ Last updated: 2021-12-21T22:42:13.000Z ## 1\. Overview In this tutorial, we're going to **select rows in Pandas DataFrame based on column values**. Selecting rows in Pandas terminology is known as indexing. We'll first look into boolean indexing, then indexing by label, the positional indexing, and finally the df.query() API. We will cover methods like `.loc` / `iloc` / `isin()` and some caveats related to their usage. ## 2\. Setup In the post, we'll use the following DataFrame, which consists of several rows and columns: ```python import pandas as pd df = pd.read_csv('https://raw.githubusercontent.com/softhints/Pandas-Tutorials/master/data/csv/extremes.csv') ``` DataFrame looks like: | Continent | Highest point | Elevation high | Lowest point | Elevation low | | ------------- | ----------------- | -------------- | ----------------- | ------------- | | Asia | Mount Everest | 8848 | Dead Sea | −427 | | South America | Aconcagua | 6960 | Laguna del Carbón | −105 | | North America | Denali | 6198 | Death Valley | −86 | | Africa | Mount Kilimanjaro | 5895 | Lake Assal | −155 | | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | ## 3\. Select rows by boolean indexing Pandas and Python use the operator `[]` for indexing. It is used for quick access in many use cases. **Note:** Since data type isn’t known in advance, directly using standard operators has some performance limits. So it's better to use optimized Pandas data access methods. ### 3.1\. What is boolean indexing? Boolean indexing in Pandas helps us to **select rows or columns by array of boolean values**. For example suppose we have the next values: `[True, False, True, False, True, False, True]` we can use it to get rows from DataFrame defined above: ```python selection = [True, False, True, False, True, False, True] df[selection] ``` result: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | ------------- | ------------- | -------------- | ------------ | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | | 2 | North America | Denali | 6198 | Death Valley | −86 | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | | 6 | Australia | Puncak Jaya | 4884 | Lake Eyre | −15 | ### 3.2\. Select Rows by Column Value with boolean indexing In most cases the boolean list will be generated by Pandas methods. For example, let's **find all rows where** the continent starts with capital A. The first step is to get a list of values where this statement is True. We are going to use string method - `str.startswith()`: ```python df['Continent'].str.startswith('A') ``` The result is Pandas Series with boolean values: ``` 0 True 1 False 2 False 3 True 4 False 5 True 6 True Name: Continent, dtype: bool ``` We can use the result (which match the DataFrame shape) for **rows selection by using operator `[]`**: ```python df[df['Continent'].str.startswith('A')] ``` only rows which match the criteria are returned: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | ---------- | ----------------- | -------------- | ------------------------- | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | | 3 | Africa | Mount Kilimanjaro | 5895 | Lake Assal | −155 | | 5 | Antarctica | Vinson Massif | 4892 | Deep Lake, Vestfold Hills | −50 | | 6 | Australia | Puncak Jaya | 4884 | Lake Eyre | −15 | ### 3.3\. Select Rows by single value **Selecting rows based on column values** is done by operator: `==`: ```python df[df['Continent'] == 'Africa'] ``` result: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | --------- | ----------------- | -------------- | ------------ | ------------- | | 3 | Africa | Mount Kilimanjaro | 5895 | Lake Assal | −155 | ### 3.4\. Select Rows by multiple value To select rows by list of values use the method - `isin()`: ```python df[df['Continent'].isin(['Africa', 'Europe'])] ``` result: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | --------- | ----------------- | -------------- | ------------ | ------------- | | 3 | Africa | Mount Kilimanjaro | 5895 | Lake Assal | −155 | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | To learn more about selection based on list of values check: [How to Select Rows by List of Values in Pandas DataFrame](https://datascientyst.com/select-rows-list-values-pandas-dataframe/) ## 4\. Select rows by property `.loc[]` Pandas has property [pandas.DataFrame.loc](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.loc.html?ref=datascientyst.com) which is documented as: > Access a group of rows and columns by label(s) or a boolean array. You can use `.loc[]` with: - a boolean array - labels The boolean indexing with `.loc[]` is similar to the point 3: ```python df.loc[df['Continent'] == 'Africa'] ``` **Note:** Note that contrary to usual Python slices, both the start and the stop are included ### 4.1\. `.loc[]` for single value Now let's cover the label usage for a single value - selecting the row with label 0: ```python df.loc[0] ``` result into Series: ``` Continent Asia Highest point Mount Everest Elevation high 8848 Lowest point Dead Sea Elevation low −427 Name: 0, dtype: object ``` ### 4.2\. `.loc[]` for multiple labels And for **selection rows by multiple labels** \- this time we pass a list with multiple values: ```python df.loc[[0,2,4]] ``` result: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | ------------- | ------------- | -------------- | ------------ | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | | 2 | North America | Denali | 6198 | Death Valley | −86 | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | In the above cases the labels are auto generated. In some cases you will have column set as index: ```python df.set_index('Continent') ``` result: | | Highest point | Elevation high | Lowest point | Elevation low | | ------------- | ------------- | -------------- | ----------------- | ------------- | | Continent | | | | | | Asia | Mount Everest | 8848 | Dead Sea | −427 | | South America | Aconcagua | 6960 | Laguna del Carbón | −105 | then you can select by labels by giving the column values: ```python df.set_index('Continent').loc["Africa"] ``` result is: ``` Highest point Mount Kilimanjaro Elevation high 5895 Lowest point Lake Assal Elevation low −155 Name: Africa, dtype: object ``` ## 5\. Select rows by positions - `iloc[]` It's possible to use the **row positions for selection - positional indexing**. This is available by property: [pandas.DataFrame.iloc](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.iloc.html?ref=datascientyst.com). For the selection we are going to: - use numpy - build a mask for highest points which contains `Mount` - get all positions of those records - select rows by their positions ```python import numpy as np mask = df['Highest point'].str.contains('Mount') pos = np.flatnonzero(mask) df.iloc[pos] ``` Only 3 rows are selected: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | --------- | ----------------- | -------------- | ------------ | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | | 3 | Africa | Mount Kilimanjaro | 5895 | Lake Assal | −155 | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | Note that Pandas has method: [pandas.DataFrame.mask](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.mask.html?ref=datascientyst.com) ## 6\. Select rows by - `df.query()` The final option which we will cover is method: [pandas.DataFrame.query](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.query.html?ref=datascientyst.com). The syntax reminds to SQL: ```python df.query('Continent == "Europe"') ``` Only one row is returned( as a DataFrame and not Series!): | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | --------- | ------------- | -------------- | ------------ | ------------- | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | For more advanced examples please check(multiple criteria, partial match spaces): [How to Select Rows by List of Values in Pandas DataFrame](https://datascientyst.com/select-rows-list-values-pandas-dataframe/) ## 6\. Conclusion Selecting rows in Pandas is available by many options and methods. Unfortunately, some of them might be slow for big DataFrames. Also, since the selection depends on the data types and criteria, choosing the optimal way is important. My recommendation is to start with the simplest one for you and search for a better solution only in case of a need. Check out the complete code for How to Select Rows by Column Value in Pandas on [GitHub](https://github.com/softhints/Pandas-Tutorials/blob/master/row/select-rows-column-value-pandas.ipynb?ref=datascientyst.com). ### How to Select Rows by List of Values in Pandas DataFrame URL: https://datascientyst.com/select-rows-list-values-pandas-dataframe/ Last updated: 2021-12-21T07:53:14.000Z ## 1\. Overview In this post, we'll explore how to **select rows by list of values in Pandas**. We'll discuss the basic indexing, which is needed in order to perform selection by values. Row selection is also known as indexing. There are several ways to select rows by multiple values: - `isin()` \- Pandas way - exact match from list of values - `df.query()` \- SQL like way - `df.loc` \+ `df.apply(lambda` \- when custom function is needed to be applied; more flexible way ## 2\. Setup We'll use the following DataFrame, which consists of several rows and columns: ```python import pandas as pd df = pd.read_csv('https://raw.githubusercontent.com/softhints/Pandas-Tutorials/master/data/csv/extremes.csv') ``` DataFrame looks like: | Continent | Highest point | Elevation high | Lowest point | Elevation low | | ------------- | ----------------- | -------------- | ----------------- | ------------- | | Asia | Mount Everest | 8848 | Dead Sea | −427 | | South America | Aconcagua | 6960 | Laguna del Carbón | −105 | | North America | Denali | 6198 | Death Valley | −86 | | Africa | Mount Kilimanjaro | 5895 | Lake Assal | −155 | | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | Then let's suppose that we would like to search by different lists from single or multiple columns: - \['America', 'Europe', 'Asia'\] - \['America', 'Europe', 'Asia'\] and range of numbers \[5000, 8000\] ## 3\. Select rows by values with method isin() The method [pandas.DataFrame.isin](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.isin.html?ref=datascientyst.com) is probably the most popular way for selection by exact match of list of values. ### 3.1\. Positive selection Let's find all rows which match exactly one of the next values in column - Continent: - \['America', 'Europe', 'Asia'\] ```python sel_continents = ['America', 'Europe', 'Asia'] df[df['Continent'].isin(sel_continents)] ``` result will be: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | --------- | ------------- | -------------- | ------------ | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | **Note:** Since method \`isin\` works by exact match rows for America are not selected. ### 3.2\. Negative selection To get all rows which doesn't include list of values you can use operator: `~`: ```python df[~df['Continent'].isin(sel_continents)] ``` result: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | ------------- | ----------------- | -------------- | ------------------------- | ------------- | | 1 | South America | Aconcagua | 6960 | Laguna del Carbón | −105 | | 2 | North America | Denali | 6198 | Death Valley | −86 | | 3 | Africa | Mount Kilimanjaro | 5895 | Lake Assal | −155 | | 5 | Antarctica | Vinson Massif | 4892 | Deep Lake, Vestfold Hills | −50 | | 6 | Australia | Puncak Jaya | 4884 | Lake Eyre | −15 | ## 4\. Query rows by values - method df.query() Let's now look at how we can do a query based on a list of values. The method: [pandas.DataFrame.query](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.query.html?ref=datascientyst.com) gives us the opportunity to select rows in SQL fashion. ### 4.1\. Query by list of values To select rows by query and multiple values use the next syntax: ```python sel_continents = ['America', 'Europe', 'Asia'] df.query('Continent in @sel_continents') ``` The resulted rows are: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | --------- | ------------- | -------------- | ------------ | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | × You can refer to variables in the environment by prefixing them with an ‘@’ character like @a + b. **Note:** For columns with spaces in their name, you can use backtick quoting: ```python df.query('B == `C C`') ``` which is equivalent to: ```python df[df.B == df['C C']] ``` ### 4.2\. Select rows by range of numbers In case of range of values `query` method can be used as follows: ```python df.query('5000 < `Elevation high` < 6000') ``` result: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | --------- | ------------- | -------------- | ------------ | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | ### 4.3\. Select rows by multiple queries Selecting by `df.query()` is considered to be the fastest way to get rows by values. So it's suitable for searching based on multiple criterias: ```python sel_continents = ['America', 'Europe', 'Asia'] df.query('(Continent in @sel_continents) and (`Elevation high` > 8000)') ``` This returns: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | --------- | ------------- | -------------- | ------------ | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | ## 5\. Select rows by values - `df.loc` \+ `df.apply(lambda` Finally let's check a slower but **more flexible way of indexing by list of values in Pandas.** ### 5.1\. Select rows by function The basic selection by `df.loc` \+ `df.apply(lambda` does boolean indexing based on lambda function applied over the rows: ```python sel_continents = ['America', 'Europe', 'Asia'] df.loc[df.apply(lambda x: x['Continent'] in sel_continents, axis=1)] ``` result: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | --------- | ------------- | -------------- | ------------ | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | × To select rows by a custom function use the next syntax: ```python def some_function(x): if x in sel_continents: return True else: return False sel_continents = ['America', 'Europe', 'Asia'] df[df['Continent'].apply(some_function)] ``` ### 5.2\. Case-insensitive search Now let's say that you would like to perform **case-insensitive search by list of values:** ```python sel_continents = ['America', 'Europe', 'Asia'] sel_continents = [item.lower() for item in sel_continents] df.loc[df.apply(lambda x: x['Continent'].lower() in sel_continents, axis=1)] ``` output: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | --------- | ------------- | -------------- | ------------ | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | ### 5.3\. Select rows by partial match based on list of values Custom functions will help us to select rows by multiple values with partial match. For example getting list of rows which contains part of the next values: - \['America', 'Europe', 'Asia'\] ```python def some_function(x): flag = False for continent in sel_continents: if continent in x or continent == x: flag = True return flag sel_continents = ['America', 'Europe', 'Asia'] df[df['Continent'].apply(some_function)] ``` in the result we have the rows for America: | | Continent | Highest point | Elevation high | Lowest point | Elevation low | | - | ------------- | ------------- | -------------- | ----------------- | ------------- | | 0 | Asia | Mount Everest | 8848 | Dead Sea | −427 | | 1 | South America | Aconcagua | 6960 | Laguna del Carbón | −105 | | 2 | North America | Denali | 6198 | Death Valley | −86 | | 4 | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | ## 6\. Conclusion In this article, we looked at different solutions for selection and indexing of rows based on values. We focused on the basics and followed some advanced techniques like SQL like query and selection by functions. The code for the examples is available: [on GitHub](https://github.com/softhints/Pandas-Tutorials/blob/master/row/select-rows-list-values-pandas-dataframe.ipynb?ref=datascientyst.com) ### How to Easily Create Dummy DataFrame with Test Data? URL: https://datascientyst.com/create-easily-dummy-dataframe-test-data/ Last updated: 2021-12-30T10:06:02.000Z ## 1\. Overview In this quick tutorial, we'll be discussing the creation of DataFrame for tests or Dummy DataFrame. We'll take a quick look at different ways to create DataFrames with data suitable for different kinds of tests. ## 2\. Fake Realistic DataFrame To create DataFrame with fake realistic data please refer to this article: [How To Make a Fake Data Set in Python and Pandas](https://datascientyst.com/make-fake-data-set-python-pandas/) ## 3\. Empty DataFrame Creation of empty DataFrame for tests is covered in: ## 4\. Dummy DataFrame with Random Numbers DataFrame creation with random numbers is covered in this post: ## 5\. Dummy DataFrame with Time Series data It's possible to create a Series or DataFrame with time series data for tests. Both have index datetime and numeric values. ### 5.1\. Series To create time series with dummy data we can use method `makeTimeSeries`: ```python import pandas as pd from pandas.util.testing import makeTimeSeries df = makeTimeSeries() df.head() ``` result: ``` 2000-01-03 -2.066624 2000-01-04 -0.897585 2000-01-05 -0.458669 2000-01-06 1.038565 2000-01-07 -0.058897 Freq: B, dtype: float64 ``` There will be 30 values created. **Note:** If you get: **FutureWarning: pandas.util.testing is deprecated** or error you can replace: \`pandas.util.testing\` for \`pandas.testing\` To learn more about the hidden methods in Pandas you can search for them in Pandas repo: [makeMissingDataframe](https://github.com/pandas-dev/pandas/search?q=makeMissingDataframe&ref=datascientyst.com) ### 5.2\. DataFrame We can create dummy DataFrame with time series data by using hidden method(not in the official docs) `makePeriodFrame`: ```python import pandas as pd from pandas.util.testing import makePeriodFrame df = makePeriodFrame() df.head() ``` result: | | A | B | C | D | | ---------- | ---------- | ---------- | ---------- | ---------- | | 2000-01-03 | 0.328536 | \-0.356004 | 0.913654 | \-0.778496 | | 2000-01-04 | \-0.501882 | \-0.476268 | \-1.426332 | 1.342977 | | 2000-01-05 | \-0.816284 | 0.028936 | \-0.890490 | \-0.531186 | | 2000-01-06 | 0.806613 | \-0.364285 | \-0.285627 | \-1.083140 | | 2000-01-07 | \-0.250598 | 0.223422 | 0.760379 | 0.220205 | **Note:** If you get: **FutureWarning: pandas.util.testing is deprecated** or error you can replace: \`pandas.util.testing\` for \`pandas.testing\` ## 6\. Dummy DataFrame with missing data What about DataFrame with missing data? Pandas has a method for that case - `makeMissingDataframe`. This method generates DataFrame with 5 columns with numeric data. There is missing data on random positionse0eb6ae0eb6a: ```python import pandas as pd from pandas.util.testing import makeMissingDataframe df = makeMissingDataframe() df.shape ``` this will result into: ``` (30, 4) ``` And data will look like: | | A | B | C | D | | ---------- | ---------- | ---------- | ---------- | ---------- | | PlqBxeS6rd | 0.777225 | NaN | 2.474689 | \-0.636386 | | SwbHSSz1EM | \-0.281544 | \-0.900471 | \-1.284081 | NaN | | d59w8ixnJO | 1.205294 | 0.170578 | \-0.930505 | \-0.095696 | | D5stwhgvIN | 0.037623 | \-1.088020 | 0.058592 | 0.371408 | | fdIzwVg1SY | NaN | \-1.294436 | 1.019611 | \-1.128139 | ## 7\. Dummy DataFrame with mixed data To create DataFrame with mixed test data use the following code: ```python import pandas as pd from pandas.util.testing import makeMixedDataFrame df = makeMixedDataFrame() df.shape ``` result: | | A | B | C | D | | - | --- | --- | ---- | ---------- | | 0 | 0.0 | 0.0 | foo1 | 2009-01-01 | | 1 | 1.0 | 1.0 | foo2 | 2009-01-02 | | 2 | 2.0 | 0.0 | foo3 | 2009-01-05 | | 3 | 3.0 | 1.0 | foo4 | 2009-01-06 | | 4 | 4.0 | 0.0 | foo5 | 2009-01-07 | ## 8\. Dummy DataFrame with periods Finally let's cover the case when we need a dataframe with date time values. Method for Pandas testing purposes `makeTimeDataFrame` will help us: ```python import pandas as pd from pandas.util.testing import makePeriodFrame df = makePeriodFrame() df.shape ``` result: | | A | B | C | D | | ---------- | ---------- | ---------- | ---------- | ---------- | | 2000-01-03 | 0.962437 | \-0.097855 | \-1.304872 | \-1.364205 | | 2000-01-04 | 0.829495 | \-1.404300 | \-0.992702 | 0.562428 | | 2000-01-05 | \-0.012089 | 1.445668 | \-0.108887 | \-1.282800 | | 2000-01-06 | 0.301585 | 0.661285 | \-0.250997 | 0.570749 | | 2000-01-07 | 0.462778 | 1.721186 | \-1.315672 | 0.201622 | If you wonder what is the difference between methods: - `makePeriodFrame()` \- `dtype='period[B]')` - `makeTimeDataFrame()` \- `dtype='datetime64[ns]', freq='B')` The type of the column is the answer. ## Conclusion So, we've taken a deep dive into creating dummy DataFrames for test purposes. And we've taken a look at some edge cases which will speed up your tests. High quality dummy data is essential for the start phase of Data Science projects. ### How To Make a Fake Data Set in Python and Pandas URL: https://datascientyst.com/make-fake-data-set-python-pandas/ Last updated: 2021-12-19T07:01:13.000Z ## 1\. Overview Do you want to **create a test data set with fake data in Python and Pandas?** Pandas in combination with Faker ease creation of test DataFrames with fake and safe for sharing data. Final output will be DataFrame which can be exported to: - CSV / Excel - JSON - XML - SQL inserts - or text file High quality test data might be crucial for the success of a given product. On the other hand, using sensitive data might cause legal issues. So the way to go is to **create high quality fake data in Pandas and Python.** ## 2\. Setup First we need install additional library - `Faker`: - [Faker - PyPI](https://pypi.org/project/Faker/?ref=datascientyst.com) - [Faker Docs](https://faker.readthedocs.io/en/master/?ref=datascientyst.com) by: ```python pip install Faker ``` This library can generate data by several providers: - [standard](https://faker.readthedocs.io/en/master/providers.html?ref=datascientyst.com) \- few examples below - bank - addresses - person - [community](https://faker.readthedocs.io/en/master/communityproviders.html?ref=datascientyst.com) - music - vehicle - [locales](https://faker.readthedocs.io/en/master/locales.html?ref=datascientyst.com) - Locale en\_US - faker.providers.address - faker.providers.automotive - Locale es - faker.providers.address - faker.providers.person ## 3\. Generate data with Faker We are going to cover the most popular use case of Faker. This step will explain different techniques in detail. To generate good data you need to apply one or several from them. ### 3.1\. Locales - country specific data Faker uses locales in order to generate country specific data like names and addresses. Generating Italian names: ```python from faker import Faker fake = Faker('it_IT') for _ in range(5): print(fake.name()) ``` result: ``` Liliana Vespa Sig.ra Vanessa Malacarne Gemma Fioravanti Vittorio Antonello Sandro Trebbi ``` You can generate data for multiple locales - Italy, US and Japan: ```python from faker import Faker fake = Faker(['it_IT', 'en_US', 'ja_JP']) for _ in range(10): print(fake.name()) ``` this will result into: ``` 高橋 聡太郎 藤井 桃子 Aurora Scaramucci Sarah Beltran Rosina Durante Sig.ra Melina Morgagni 佐藤 修平 林 修平 Charles Marks Lazzaro Ammaniti ``` ### 3.2\. Dynamic Providers - custom data Faker offers an elegant way of generating data from a custom set of elements. Let say that you like to generate fake skills from next set of words: - "Python" - "Pandas" - "Linux" **Dynamic providers** are the Faker way to **generate custom data**: ```python skill_provider = DynamicProvider( provider_name="skills", elements=["Python", "Pandas", "Linux", "SQL", "Data Mining"], ) fake = Faker('en_US') fake.add_provider(skill_provider) fake.skills() ``` ### 3.3\. Numbers and ranges Generating numbers and ranges is fine with pure Python. The below examples show how to generate random integer between: 0 and 15: ```python random.randint(0,15) ``` and random numbers in range with steps: ```python random.randrange(75000,150000, 5000) ``` ### 3.4\. Dates For dates we can use the Faker methods - `date()`. It accepts format, start and end time: ```python fake.date(pattern="%Y-%m-%d", end_datetime=datetime.date(1995, 1,1)) ``` ### 3.5\. Correlated Data Generation of correlated data is very important for test quality. What does it mean correlated fake data? Generating pairs of country and city should be correlated: - US and Paris - not correlated - US and NY - correlated The same is for: - names and email - birth date and experience - nationality and native language - experience and salary #### Locales In most cases custom solutions are required in order to get meaningful data. For example by using locales: ```python from faker import Faker en_us_faker = Faker('en_US') it_it_fake = Faker('it_IT') print(f'{en_us_faker.city()}, USA') print(f'{it_it_fake.city()}, Italy') ``` result would be: - East Jonathanbury, USA - Biagio calabro, Italy #### Customization Another option is by using customization. They are described here: [Faker Customization](https://github.com/faker-ruby/faker?ref=datascientyst.com#customization) ## 4\. Define the output list of fields Next we should define what fields we need for the final data set. In this article we are going to generate personal data - employees with names, skills and salaries. The list of fields for the final data set is: - first name - last name - birth date - email - experience - start year - salary - main skill - nationality - city Add description for each field if needed like: - first name - Italian and American names - birth date - dates between 1990 and 2000 ## 5\. Create DataFrame with Fake Data Finally let's check how to **generate the fake data and then stored it to a DataFrame**. The code below will generate 50 records of personal fake data: ```python import csv import pandas as pd from faker import Faker import datetime import random from faker.providers import DynamicProvider skill_provider = DynamicProvider( provider_name="skills", elements=["Python", "Pandas", "Linux", "SQL", "Data Mining"], ) def fake_data_generation(records): fake = Faker('en_US') employee = [] fake.add_provider(skill_provider) for i in range(records): first_name = fake.first_name() last_name = fake.last_name() employee.append({ "First Name": first_name, "Last Name": last_name, "Birth Date" : fake.date(pattern="%Y-%m-%d", end_datetime=datetime.date(1995, 1,1)), "Email": str.lower(f"{first_name}.{last_name}@fake_domain-2.com"), "Hobby": fake.word(), "Experience" : random.randint(0,15), "Start Year": fake.year(), "Salary": random.randrange(75000,150000, 5000), "City" : fake.city(), "Nationality" : fake.country(), "Skill": fake.skills() }) return employee df = pd.DataFrame(fake_data_generation(50)) ``` result(output is transposed for redness purpose): | | 0 | 1 | | ----------- | ------------------------------------ | -------------------------------- | | First Name | Christopher | April | | Last Name | Davis | Collins | | Birth Date | 1977-09-14 | 1974-06-08 | | Email | christopher.davis@fake\_domain-2.com | april.collins@fake\_domain-2.com | | Hobby | wear | family | | Experience | 0 | 6 | | Start Year | 1991 | 1994 | | Salary | 100000 | 85000 | | City | Charleschester | East Josephchester | | Nationality | Panama | Cambodia | | Skill | Pandas | Python | Now the DataFrame can be easily exported to CSV - `to_csv` etc. The final result is visible on the image below: ![](https://datascientyst.com/content/images/2021/12/make-fake-data-set-python-pandas.png) ## 6\. Create big CSV file with Fake Data As a bonus we can see how to\*\* generate huge data sets with fake data\*\*. This time we are going to write directly to a CSV file for performance sake: ```python import csv from faker import Faker import datetime import random from faker.providers import DynamicProvider skill_provider = DynamicProvider( provider_name="skills", elements=["Python", "Pandas", "Linux", "SQL", "Data Mining"], ) def fake_data_generation(records, headers): fake = Faker('en_US') fake.add_provider(skill_provider) with open("employee.csv", 'wt') as csvFile: writer = csv.DictWriter(csvFile, fieldnames=headers) writer.writeheader() for i in range(records): first_name = fake.first_name() last_name = fake.last_name() print({ "First Name": first_name, "Last Name": last_name, "Birth Date" : fake.date(pattern="%Y-%m-%d", end_datetime=datetime.date(1995, 1,1)), "Email": str.lower(f"{first_name}.{last_name}@fake_domain-2.com"), "Hobby": fake.word(), "Experience" : random.randint(0,15), "Start Year": fake.year(), "Salary": random.randrange(75000,150000, 5000), "City" : fake.city(), "Nationality" : fake.country(), "Skill": fake.skills() }) number_records = 100 fields = ["First Name", "Last Name", "Birth Date", "Email", "Hobby", "Experience", "Start Year", "Salary", "City", "Nationality", "Skill"] fake_data_generation(number_records, fields) ``` The output will be `employee.csv` in the current working folder. We providing the headers and the number of the records: ```python number_records = 100 fields = ["First Name", "Last Name", "Birth Date", "Email", "Hobby", "Experience", "Start Year", "Salary", "City", "Nationality", "Skill"] ``` ## 7\. Conclusion So, we've taken a look into fake data generation; exploring the basics and its basic use. We've taken a look at some details like correlated data and quality fake data. Finally, we saw how to export the saved data to different formats and optimized the process. The code is [available on GitHub](https://github.com/softhints/Pandas-Tutorials/blob/master/create/make-fake-data-set-python-pandas.ipynb?ref=datascientyst.com). ### How to Create Empty DataFrame in Pandas URL: https://datascientyst.com/create-empty-dataframe-pandas/ Last updated: 2021-12-17T22:33:31.000Z ## 1\. Overview This short article describes **how to create empty DataFrame in Pandas.** Generally, there are three options to create empty DataFrame: - Empty DataFrame - Empty DataFrame with columns - Empty DataFrame with index and columns In general it's not recommended to create empty DataFrame and append rows to it. ## 2\. Create Empty DataFrame Let's start by creating completely empty DataFrame: - no data - no index - no columns The constructor called without any parameters as: ```python import pandas as pd df = pd.DataFrame() ``` Will create a DataFrame object that doesn't have data and attributes. If you print it, you will get: ``` Empty DataFrame Columns: [] Index: [] ``` To append rows and add columns to empty DataFrame - you can use bracket notation and assign values to it. So in brackets we have the column name and the lists contains the row values: ```python df['Rank'] = ['silver', 'gold', 'gold'] df['score'] = [54, 75, 87] ``` Now the result will be: | | Rank | score | | - | ------ | ----- | | 0 | silver | 54 | | 1 | gold | 75 | | 2 | gold | 87 | ## 3\. Create Empty DataFrame with index and columns What if you like to make **DataFrame without data - only with index and columns**. The DataFrame constructor needs two parameters: `columns` and `index`: ```python import pandas as pd df = pd.DataFrame(columns = ['Score', 'Rank'], index = ['James', 'Jim']) ``` our DataFrame without any data looks like: | | Score | Rank | | ----- | ----- | ---- | | James | NaN | NaN | | Jim | NaN | NaN | To set data for any of the existing rows we can use property `loc` which is described as: > Access a group of rows and columns by label(s) or a boolean array. To set values for Jim we can do: ```python df.loc['Jim'] = [50, 'gold'] ``` result: | | Score | Rank | | ----- | ----- | ---- | | James | NaN | NaN | | Jim | 50 | gold | ## 4\. Create Empty DataFrame with column names To **create a DataFrame which has only column names** we can use the parameter `column`. We are using the DataFrame constructor to create two columns: ```python import pandas as pd df = pd.DataFrame(columns = ['Score', 'Rank']) print(df) ``` result: ``` Empty DataFrame Columns: [Score, Rank] Index: [] ``` If you display it you will get: | | Score | Rank | | | ----- | ---- | To append rows to this DataFrame we can use method [append()](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.append.html?ref=datascientyst.com) and provide dictionary with values: ```python df.append({'Score' : '68', 'Rank' : 'gold'}, ignore_index = True) ``` Result is: | | Score | Rank | | - | ----- | ---- | | 0 | 68 | gold | Parameter `ignore_index` is used in order to avoid error: > TypeError: Can only append a dict if ignore\_index=True ## 5\. Efficient way to append rows to DataFrame And the way above is not very efficient for multiple rows. A better approach to append rows to empty DataFrame is by method [concat()](https://pandas.pydata.org/docs/reference/api/pandas.concat.html?ref=datascientyst.com#pandas.concat): So in order to append multiple rows to empty or full DataFrame `df` like: | | Score | Rank | | ----- | ----- | ---- | | James | NaN | NaN | | Jim | 50 | gold | is by concatenating the two DataFrames: ```python df_a = pd.DataFrame({'Score' : ['68', '5'], 'Rank' : ['gold', 'zero']}, index=['Joe', 'Jery']) pd.concat([df, df_a]) ``` result: | | Score | Rank | | ----- | ----- | ---- | | James | NaN | NaN | | Jim | 50 | gold | | Joe | 68 | gold | | Jery | 5 | zero | ## 6\. Conclusion And now we're able to create empty DataFrame in several ways. We looked at how to add data, index and columns. We covered more efficient way of doing it. Now you know more about DataFrame and how you can add rows and data to it. ### How to Iterate Over Rows in Pandas DataFrame URL: https://datascientyst.com/iterate-over-rows-pandas-dataframe/ Last updated: 2021-12-11T20:00:32.000Z ## 1\. Overview In this quick guide, we're going to see **how to iterate over rows in Pandas DataFrame**. Pandas offer several different methods for iterating over rows like: - [DataFrame.iterrows()](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.iterrows.html?ref=datascientyst.com#pandas-dataframe-iterrows) - [DataFrame.itertuples()](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.itertuples.html?ref=datascientyst.com) This article will explain the most common ways. **Note:** Have in mind that iterating over rows is pretty slow operation and not needed in most cases. There is even warning on [Pandas docs](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/basics.html?ref=datascientyst.com#iteration): > Iterating through pandas objects is generally slow. ## 2\. Setup In the article, we'll use the small DataFrame, which consists of several rows: ```python import pandas as pd df = pd.read_csv('https://raw.githubusercontent.com/softhints/Pandas-Tutorials/master/data/csv/extremes.csv') ``` DataFrame looks like: | Continent | Highest point | Elevation high | Lowest point | Elevation low | | ------------- | ----------------- | -------------- | ----------------- | ------------- | | Asia | Mount Everest | 8848 | Dead Sea | −427 | | South America | Aconcagua | 6960 | Laguna del Carbón | −105 | | North America | Denali | 6198 | Death Valley | −86 | | Africa | Mount Kilimanjaro | 5895 | Lake Assal | −155 | | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | ## 3\. Iterate Using `iterrows()` Let's start by method `iterrows()` from the DataFrame class which iterates over rows and returns pairs of (index, Series). This is the most **popular way for iteration in Pandas DataFrame**: ```python for index, row in df.iterrows(): print(index, row) ``` This will result into index of the first row and all values, then the second etc: ``` 0 Continent Asia Highest point Mount Everest Elevation high 8848 Lowest point Dead Sea Elevation low −427 Name: 0, dtype: object 1 Continent South America Highest point Aconcagua Elevation high 6960 Lowest point Laguna del Carbón Elevation low −105 Name: 1, dtype: object ``` Row values are accessible with bracket notation: `row['Continent']` ```python for index, row in df.iterrows(): print(index, row['Continent'], row['Elevation high']) ``` The output is: ``` 0 Asia 8848 1 South America 6960 2 North America 6198 3 Africa 5895 4 Europe 5642 5 Antarctica 4892 6 Australia 4884 ``` The image below demonstrates how the method works: ![](https://datascientyst.com/content/images/2021/12/iterate-over-rows-pandas-dataframe.png) ## 4\. Using `df.itertuples()` Another method which iterates over rows is: `df.itertuples()`. **`df.itertuples` is a faster for iteration over rows in Pandas.** To loop over all rows in a DataFrame by `itertuples()` use the next syntax: ```python for row in df.itertuples(): print(row) ``` this will result into(all rows are returned as namedtuples): ``` Pandas(Index=0, Continent='Asia', _2='Mount Everest', _3=8848, _4='Dead Sea', _5='−427') Pandas(Index=1, Continent='South America', _2='Aconcagua', _3=6960, _4='Laguna del Carbón', _5='−105') ``` In order to access only the first row we need to use `next(iter(` \- because generator is returned by this method: ```python next(iter(df.itertuples(index=True, name='Point'))) ``` output: ``` Point(Index=0, Continent='Asia', _2='Mount Everest', _3=8848, _4='Dead Sea', _5='−427') ``` **Note:** itertuples() have the parameter \`name\`. If it's missing then the default value is Pandas. Accessing row data with `itertuples()` is available by indices - integers or slices: ```python for row in df.itertuples(index=True, name='Point'): print(row[3], row[2]) ``` rows are returned as: ``` 8848 Mount Everest 6960 Aconcagua 6198 Denali ``` **Note:** namedtuples are subclasses of tuples. You can think of them like something between dict and tuple. It adds more features in comparison to tuples. Check more for namedtuples: [collections.namedtuple](https://docs.python.org/3/library/collections.html?ref=datascientyst.com#collections.namedtuple) ## 5\. Faster Iteration over rows For bigger datasets a faster solution is required. There are many options available if you need to speed up the loop over rows. For example you can use frameworks like: - [Dask](https://dask.org/?ref=datascientyst.com) - [Modin](https://modin.readthedocs.io/en/stable/?ref=datascientyst.com) and others like: Vaex, Ray, RAPIDS Best for pure Pandas is to use vectorization for your operations. Another option for processing all rows is list comprehensions. In the example below you can iterate over each row and get values for 2 columns: ```python [print(x, y) for x, y in zip(df['Continent'], df['Highest point'])] ``` result: ``` Asia Mount Everest South America Aconcagua North America Denali ``` or applying some function: ```python def func(x, y): return x + ' : ' + y result = [func(x, y) for x, y in zip(df['Continent'], df['Highest point'])] result ``` result: ``` ['Asia : Mount Everest', 'South America : Aconcagua', ... 'Antarctica : Vinson Massif', 'Australia : Puncak Jaya'] ``` ## 6\. Conclusion In this post, we looked at different ways for iterating over rows in Pandas. We focused on basic functionality, but also compared the advantages of different methods. In general for beginners and medium datasets - `df.iterrows()` is the way to go. For bigger datasets and more advanced users - `itertuples()`, list comprehension or custom solution are better. As usual the code examples are available on [GitHub](https://github.com/softhints/Pandas-Tutorials/blob/master/row/iterate-over-rows-pandas-dataframe.ipynb?ref=datascientyst.com). ### What is Pandas Series URL: https://datascientyst.com/what-is-pandas-series/ Last updated: 2022-12-08T08:03:18.000Z In this guide, we'll learn about one of the two main data structures in **Pandas - Series**. Our goal will be to understand the usage of this structure. Understanding what a Pandas Series will avoid simple mistakes in future. Pandas Series play a major role in data wrangling and transformation. ## Step 1: What is Pandas Series There two main data structures in Pandas: - [Series](https://pandas.pydata.org/docs/reference/api/pandas.Series.html?ref=datascientyst.com) - [DataFrames](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.html?ref=datascientyst.com) The official documentation describes Series like: > One-dimensional ndarray with axis labels (including time series). So Series is a one-dimensional array. It has labels to access data. You can imagine it like a sequence of post boxes - each has an address and can store different items. **Note:** What is ndarray? \`ndarray\` is a multi-dimensional array in Numpy. It should have homogeneous and fixed-size items. The image below illustrates the Series visually. There are two main parts: - labels (also index or axis=0) - it can be set explicitly ot auto generated - data (also - values) - can store different types of data and even empty values like null, None, NA ![pandas-series-index-values](https://datascientyst.com/content/images/2021/12/pandas-series-index-values.png) **Note:** Labels must be a hashable type (no need to be unique). Series is the building block for DataFrames. ## Step 2: Create a Pandas Series In order to create Pandas Series we can use the following constructor: ```python class pandas.Series(data=None, index=None, dtype=None, name=None, copy=False, fastpath=False) ``` **Below you can find example of creating Series from a dict**: ```python import pandas as pd d = {'x': 1, 'y': 2, 'z': 3} s = pd.Series(data=d, index=['x', 'y', 'z']) ``` In the example above we have labeled data: `d = {'x': 1, 'y': 2, 'z': 3}` and index(which should match): `index=['x', 'y', 'z']` This would result into: ``` x 1 y 2 z 3 dtype: int64 ``` If the index is skipped: ```python d = {'x': 1, 'y': 2, 'z': 3} s = pd.Series(data=d) ``` result would be the same: ``` x 1 y 2 z 3 dtype: int64 ``` and if we change order of the index: ```python d = {'x': 1, 'y': 2, 'z': 3} s = pd.Series(data=d, index=['x', 'z', 'y']) ``` we will get different order in the Series: ``` x 1 z 3 y 2 dtype: int64 ``` **Note:** If data is dict-like and index is None, then the keys in the data are used as the index. If the index is not None, the resulting Series is reindexed with the index values. As you can see the Series has a dtype. Dtype represents the type of the stored data. ![](https://datascientyst.com/content/images/2021/12/pandas-series.png) Since in the above we store integers Pandas creates the Series with `dtype: int64`. ## Step 3: Create Pandas Series from list We can create **Pandas Series also by providing iterable like a list**: ```python s = pd.Series(['a', 'b', 'c']) ``` This time the dtype is object: ``` 0 a 1 b 2 c dtype: object ``` If index is not provided as in the example with the dict - then automatic labels will be applied on the Series starting from 0. ## Step 4: Pandas Series Attributes/Properties **Attributes pr properties of Series** are storing important information or performing features. Attributes can be invoked in this way: ```python s.dtype ``` which will result into: `dtype('int64')` Below you can find some of the most used Series attributes. Let have the next Series: ``` x 1 y 2 z 3 dtype: int64 ``` Below you can find the attribute, the explanation and the result of the execution. - **dtype** > Return the dtype object of the underlying data. Example result: `dtype('int64')` - **index** > The index (axis labels) of the Series. Example result: `Index(['x', 'y', 'z'], dtype='object')` - **values** > Return Series as ndarray or ndarray-like depending on the dtype. Example result: `array([1, 2, 3])` - **shape** > Return a tuple of the shape of the underlying data. Example result: `(3,)` - **loc** > Access a group of rows and columns by label(s) or a boolean array. ```python s.loc[[1]] ``` Example result: ``` 1 b dtype: object ``` ## Step 5: Series Methods in Pandas The core functionality of Pandas is available via methods. In this section you can find some of the **most popular Series methods**. Currently the official documentation shows around 200 Pandas Series methods! Usually the difference between attributes and methods is the brackets - `()` Let use the next Series and check the methods: ``` x 1 y 2 z 3 dtype: int64 ``` Example usage of Series methods: ```python s.sum() ``` - **sum** > Return the sum of the values over the requested axis. Example result: `6` - **head(\[n\])** Similar one is `tail()` which returns the last n rows. > Return the first n rows. `s.head(2)`: result: ``` x 1 y 2 dtype: int64 ``` - **isna()** > Detect missing values. Example result: ``` x False z False y False dtype: bool ``` - **duplicated(\[keep\])** > Indicate duplicate Series values. Example result: ``` x False z False y False dtype: bool ``` **fillna(\[value, method, axis, inplace, ...\])** > Fill NA/NaN values using the specified method. - **sort\_values(\[axis, ascending, inplace, ...\])** > Sort by the values. For example: ```python s.sort_values(ascending=False) ``` will return sorted Series: z 3 y 2 x 1 dtype: int64 **Note:** The original Series will not be changed. The copy of it will be returned. ## Step 6: Alias Methods of Pandas Series Those are special aliases for methods like String or Datetime. The idea is to use methods from spaces like: `StringMethods`. Example usage is: ```python s.str.count('a') ``` ### Pandas Series .str > alias of pandas.core.strings.accessor.StringMethods Some of the StringMethods are: - [pandas.Series.str.replace](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.replace.html?ref=datascientyst.com) - [pandas.Series.str.contains](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.contains.html?ref=datascientyst.com) - [pandas.Series.str.split](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.split.html?ref=datascientyst.com) - [pandas.Series.str.count](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.count.html?ref=datascientyst.com) Example of using `.str` accessor in real life: [How to Replace Regex Groups in Pandas](https://datascientyst.com/replace-regex-groups-in-pandas/) ### Pandas Series .dt > alias of pandas.core.indexes.accessors.CombinedDatetimelikeProperties Some properties are: - [pandas.Series.dt.quarter](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.quarter.html?ref=datascientyst.com) - [pandas.Series.dt.to\_period](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.to%5Fperiod.html?ref=datascientyst.com) - [pandas.Series.dt.weekday](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.weekday.html?ref=datascientyst.com) Example of using `.dt` accessor in real life: [How to Extract Month and Year from DateTime column in Pandas](https://datascientyst.com/extract-month-and-year-datetime-column-in-pandas/) **Note:** You can use those aliases only if the data has a specific type. Otherwise you will get an error: For example accessing `.str` on integer values will raise `AttributeError`: > AttributeError: Can only use .str accessor with string values! ### Data Science with Python and Pandas URL: https://datascientyst.com/learn-data-science-python-pandas/ Last updated: 2021-12-08T15:40:30.000Z ## Why the Python Ecosystem and Pandas for Data Science? One of the main goals of Python has always been to ease the learning curve while remaining intuitive and powerful. The language is open source which helps the thrive of the scientific packages like: - TensorFlow - Scikit-Learn - Numpy - Keras - SciPy - Pandas The ecosystem has great support from big companies and individuals. The flat learning curve allows scientists from different areas to enter the Data Science world. Pandas sits on top of Python and Numpy and simplifies data manipulation. Pandas offer great range of functions like: - import and export of various formats - data wrangling - data cleaning - text processing - time series and much more All this makes **Pandas/Python a natural choice for learning and mastering Data Science**. ## History of Pandas and Python Python was created in the late 1980s by Guido van Rossum. The initial idea of Guido was to create language which is close to plain English, powerful, open for every one and suitable for everyday tasks. You can see the Hello world! example in python: ```python print('Hello, world!') ``` Decades later the language is in the top of the most used, loved and wanted languages: - [tiobe](https://www.tiobe.com/tiobe-index/?ref=datascientyst.com) - [PYPL](https://pypl.github.io/PYPL.html?ref=datascientyst.com) - [Stack Overflow survey](https://insights.stackoverflow.com/survey/2021?ref=datascientyst.com) Pandas was started in 2008 and became open source in 2009\. The main idea behind Pandas was to: > be the fundamental high-level building block for doing practical, real world data analysis in Python. Additionally, it has the broader goal of becoming the most powerful and flexible open source data analysis / manipulation tool available in any language. Source: [About Pandas](https://pandas.pydata.org/about/?ref=datascientyst.com) One of the most used code samples or **Hello world! in Pandas** is as simple as: ```python import pandas as pd pd.read_csv("foo.csv") ``` This single line will give you shortcut to (plus few more lines): - reshaping and pivoting - slicing, fancy indexing, and subsetting - data alignment - data cleaning - data analysis - data mining ### Learn Data Science with Data Scientyst URL: https://datascientyst.com/learn-data-science/ Last updated: 2021-12-08T15:39:53.000Z ## Learn Data Science in simple steps The website is organized in different modules based on smart organization . The guides go over the basics and have many practical examples. Each module starts with basics of the subject. In the next tutorials are covered advanced problems which can not be solved by reading the basic documentation. The goal is the reader to advance in small steps and learn all needed knowledge before diving into harder problems. The materials are based on research about what is used in Data Science. Common problems and pitfalls of Pandas are collected and presented in the course. ## Learn by practice The best way for me of learning is to: - learn in small chunks - apply the new material - think on hard problems Data Science is all about analyzing and solving hard problems. While the tool is important, the mindset of the data scientist is the core. I believe that Pandas will help you faster to enter the Data Science field while working on real world problems. The fuel of the 21st century is data which is another area where Python is great - collecting and scraping data from various sources. ## Modules and Activities The home page is organized in a way to separate and group: - Pandas functionality - Data Science techniques - Activities Pandas functionality cover many areas and methods like: - sort\_values - value\_counts - filter - drop\_duplicates Data Science techniques are present in several sections like: - text processing - time series - data cleaning Finally the site has activities like: - challenges - exercise - extra learning materials like cheat-sheets ## Open and free All materials are open for everyone and free for use. There's no plan to restrict the access to the materials. Good luck! ### How to Drop Column in Pandas URL: https://datascientyst.com/drop-column-pandas-dataframe/ Last updated: 2022-03-24T07:41:14.000Z In this quick tutorial, we will see how to **drop single or multiple columns by name or index in Pandas**. We'll first look into using the `drop()` method to: - **drop a single column** - then by using alternatives like - `del` and `df.pop` - drop column with NaN values - **finally how to drop multiple columns**. ## Setup In the post, we'll use the following DataFrame, which consists of several rows and columns: ```python import pandas as pd df = pd.read_csv('https://raw.githubusercontent.com/softhints/Pandas-Tutorials/master/data/csv/extremes.csv') ``` DataFrame looks like: | Continent | Highest point | Elevation high | Lowest point | Elevation low | | ------------- | ----------------- | -------------- | ----------------- | ------------- | | Asia | Mount Everest | 8848 | Dead Sea | −427 | | South America | Aconcagua | 6960 | Laguna del Carbón | −105 | | North America | Denali | 6198 | Death Valley | −86 | | Africa | Mount Kilimanjaro | 5895 | Lake Assal | −155 | | Europe | Mount Elbrus | 5642 | Caspian Sea | −28 | ## Step 1: Drop column by name in Pandas Let's start by using the DataFrame method `drop()` to remove a single column. **To drop column named** \- 'Lowest point' we can use the next syntax: ```python df = df.drop('Lowest point', axis=1) ``` or the equivalent: ```python df = df.drop(columns='Lowest point') ``` By default method `drop()` will return a copy. If you like to do the operation in place you can use the syntax above or parameter: ```python df.drop('Lowest point', axis=1, inplace=True) ``` **Note:** Note that method works on both axes - \`axis=1\` - means columns. After the operation the DataFrame will look like: | Continent | Highest point | Elevation high | Elevation low | | ------------- | ----------------- | -------------- | ------------- | | Asia | Mount Everest | 8848 | −427 | | South America | Aconcagua | 6960 | −105 | | North America | Denali | 6198 | −86 | | Africa | Mount Kilimanjaro | 5895 | −155 | | Europe | Mount Elbrus | 5642 | −28 | ## Step 2: Drop column by index in Pandas **To drop a column by index** we will combine: - `df.columns` - `drop()` This step is based on the previous step plus getting the name of the columns by index. So to get the first column we have: ```python df.columns[0] ``` the result is: ``` Continent ``` So to **drop the column on index 0** we can use the following syntax: ```python df.drop(df.columns[0], axis=1) ``` ## Step 3\. Drop multiple columns by name in Pandas Next let's see how to **drop multiple columns in Pandas** \- for example: "Elevation high" and "Elevation low". Again we are going to use method `drop()` by providing list of columns: ```python df.drop(["Elevation high", "Elevation low"], axis=1) ``` result: | Continent | Highest point | | ------------- | ----------------- | | Asia | Mount Everest | | South America | Aconcagua | | North America | Denali | | Africa | Mount Kilimanjaro | | Europe | Mount Elbrus | This is possible because parameter `labels` can be single or list-like. **Note:** Instead of using axis - \`labels, axis=1\` you can use parameter \`columns\`: ```python df.drop(columns=["Highest point"]) ``` ## Step 4\. Drop multiple columns by index To drop multiple columns by index we can use syntax like: ```python cols = [0, 2] df.drop(df.columns[cols], axis=1, inplace=True) ``` This will drop the first and the third column from the DataFrame ## Step 5\. Drop column with NaN in Pandas To drop column or columns which contain NaN values we can use method `dropna()`: ```python df.dropna(axis=1, how='all') ``` The parameter `how='all'` will drop all columns which contain only NaN values. **Note:** that dropna() doesn't change DataFrame in place. We need to use parameter - inplace=True to do so To drop columns with NaN values by method `dropna()` we need the following parameters: - `axis=1` \- for columns - `how` - `any` \- If any NA values are present, drop that row or column - `all` \- If all values are NA, drop that row or column - `subset` \- Labels along other axis to consider, e.g. if you are dropping rows these would be a list of columns to include ## Step 6\. Drop column with `del` and `df.pop` An alternative solution **to remove column from DataFrame** is using the Python keyword - `del`: ```python del df["Lowest point"] ``` **Note:** Note that this is going to delete the column in place. One more way to achieve the same behavior is by using method `df.pop`: ```python df.pop('Highest point') ``` This method will return the column as series: 0 Mount Everest 1 Aconcagua 2 Denali 3 Mount Kilimanjaro 4 Mount Elbrus 5 Vinson Massif 6 Puncak Jaya Name: Highest point, dtype: object At the same time will remove the column from the DataFrame. ## Conclusion & Resources In this article, we looked at different ways to drop columns in Pandas. We saw how to drop single or multiple columns. How to drop columns by index or name. How to drop columns with NaN values. We covered alternative ways for dropping columns. Finally we saw which is the most efficient way of doing it. The code for the examples is available over on GitHub in a [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/column/drop-column-from-pandas-dataframe.ipynb?ref=datascientyst.com). - [Column selection, addition, deletion](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/dsintro.html?ref=datascientyst.com#column-selection-addition-deletion) - [pandas.DataFrame.drop](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.drop.html?ref=datascientyst.com) - [pandas.DataFrame.pop](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.pop.html?ref=datascientyst.com) ### How to Install Python and Pandas on Linux URL: https://datascientyst.com/install-python-pandas-linux/ Last updated: 2022-02-01T21:20:18.000Z ## 1\. Overview In this tutorial, we'll introduce **different methods for installing Pandas and Python on Linux**. Then, we'll briefly compare the methods. Finally, we'll show how to manage multiple Python/Pandas versions on Linux. As a prerequisite to each method, we need - minimal ternimal knowledge - non-root user with sudo privileges The instructions described below have been tested on Linux Mint 19 and 20. ## 2\. Install Python on Linux / Ubuntu Python is a widely-used general-purpose, powerful, mature and high-level programming language. It is easy to learn and has a huge community. Python is one of the most liked and wanted languages according to: [stackoverflow - Python is the most wanted language for its fifth-year](https://insights.stackoverflow.com/survey/2021?ref=datascientyst.com#most-loved-dreaded-and-wanted-language-love-dread) Most Linux systems have pre-installed Python on their machine. You can check your Linux by simple command: ```python python --version python3 --version ``` result: ``` Python 2.7.17 Python 3.6.9 ``` In 2021 Python 3 is the only one which needs to be used - Python 2 was deprecated. ### 2.1\. Install Python by apt-get Most of the Linux distros offer Python packages which can be installed simply by: ```bash sudo apt-get install python3.8 ``` Confirm the disk space prompt and wait for the installation To search for other Python packages and versions you can use: ```bash apt-cache search python3 ``` or filtering the results: `apt-cache search python3 | grep Interactive`: ``` python3.6-venv - Interactive high-level object-oriented language (pyvenv binary, version 3.6) python3.7 - Interactive high-level object-oriented language (version 3.7) python3.7-venv - Interactive high-level object-oriented language (pyvenv binary, version 3.7) ``` ### 2.2\. Download and install latest Python from python.org If we want to use the latest and greatest version of Python, often manual installation is the way to go. This means downloading the package from the [Python site](https://www.python.org/downloads/source/?ref=datascientyst.com). Once the selected version is downloaded for example: `Python-3.10.1.tar.xz` you can extract it by: ```bash tar -xf Python-3.10.1.tar.xz ``` Next you need to open the extracted folder: ```bash cd Python-3.* ``` Start the configuration by: ```python ./configure ``` Finally install Python by: ```python sudo make altinstall ``` **Note:** Important note: In order to prevent damage on your Linux system use \`altinstall\` instead of \`install\`. ### 2.3\. Verify the installation Finally the installation can be verified by next commands: `python3.8` or `which python3.8` result: ``` /usr/local/bin/python3.8 ``` ### 2.4\. Create a virtual environment (optional) Python offers a powerful package system `venv` which helps separate different Python packages. In simple words, you can create several virtual environments in multiple Pandas versions: - pandas 1.3.4 - pandas 1.0.0 To create new virtual environment called `pandas1`: - create folder for your virtual environments ( or select existing one) - Run command: `python3.8 -m venv pandas1` - activate the environment by: ```bash cd pandas1 source bin/activate ``` Once environment is activated you will see change in the terminal: ```python (pandas1) $ deactivate ``` The command above deactivates the environment. ## 3\. Install Pandas on Linux ### 3.1\. Install Pandas by Pypi Next step is to install Pandas. The most popular way of installing Pandas is by running: ```python pip install pandas ``` You can find more information for Pandas on: [pandas - pypi.org](https://pypi.org/project/pandas/?ref=datascientyst.com). ### 3.2\. Install Pandas by Anaconda If you like to use alternative installation methods you can check the official docs: [Installation](https://pandas.pydata.org/pandas-docs/stable/getting%5Fstarted/install.html?ref=datascientyst.com). For example Pandas is part of [Anaconda](https://docs.continuum.io/anaconda/?ref=datascientyst.com) \- so if you install Anaconda on your system you will get Pandas: - download the [Anaconda installer for Linux](https://www.anaconda.com/download/?ref=datascientyst.com#linux) - Verify data integrity with SHA-256\. (optional but highly RECOMMENDED step) - Install Anaconda: ```bash bash ~/Downloads/Anaconda3-2020.02-Linux-x86_64.sh ``` follow the instructions or check for more details: [Installing Anaconda on Linux](https://docs.continuum.io/anaconda/install/linux/?ref=datascientyst.com) ### 3.3\. Verify Pandas installation Finally you can test Pandas installation by running next commands: ```bash pip freeze | grep pandas ``` result will be: ```bash pandas==1.3.2 ``` ## 4\. Conclusion To summarize, in this article, we've seen examples of installing Python and Pandas from a PPA and manually. We've briefly explained these installation methods. And finally, we've seen how to manage multiple Python/Pandas installations on Ubuntu systems with different package versions. ### 424-display URL: https://datascientyst.com/display/ Last updated: 2021-11-30T13:46:23.000Z Display ### 430-beginners URL: https://datascientyst.com/beginners/ Last updated: 2021-11-30T13:46:23.000Z Beginners ### 421-styling URL: https://datascientyst.com/styling/ Last updated: 2021-11-30T13:46:22.000Z Styling ### 422-table URL: https://datascientyst.com/table/ Last updated: 2021-11-30T13:46:22.000Z Table ### 415-time-series URL: https://datascientyst.com/time-series/ Last updated: 2021-11-30T13:46:21.000Z Time Series ### 420-basic-concepts URL: https://datascientyst.com/basic-concepts-11/ Last updated: 2021-11-30T13:46:21.000Z Basic concepts ### 414-duplicate URL: https://datascientyst.com/duplicate/ Last updated: 2022-11-14T15:55:13.000Z Duplicate ### 411-data-validation URL: https://datascientyst.com/data-validation/ Last updated: 2021-11-30T13:46:20.000Z Data Validation ### 412-data-cleaning URL: https://datascientyst.com/data-cleaning/ Last updated: 2021-11-30T13:46:20.000Z Data Cleaning ### 339-exercise URL: https://datascientyst.com/exercise-9/ Last updated: 2021-11-30T13:46:19.000Z Exercise ### 410-basic-concepts URL: https://datascientyst.com/basic-concepts-10/ Last updated: 2021-11-30T13:46:19.000Z Basic concepts ### 333-concat URL: https://datascientyst.com/concat/ Last updated: 2023-02-18T09:04:16.000Z concat() ### 330-basic-concepts URL: https://datascientyst.com/basic-concepts-9/ Last updated: 2021-11-30T13:46:18.000Z Basic concepts ### 331-merge URL: https://datascientyst.com/merge/ Last updated: 2021-11-30T13:46:18.000Z merge() ### 332-join URL: https://datascientyst.com/join/ Last updated: 2023-02-18T09:04:35.000Z join() ### 326-other URL: https://datascientyst.com/other/ Last updated: 2021-11-30T13:46:17.000Z Other ### 329-exercise URL: https://datascientyst.com/exercise-8/ Last updated: 2021-11-30T13:46:17.000Z Exercise ### 323-convert URL: https://datascientyst.com/convert/ Last updated: 2021-11-30T13:46:16.000Z Convert ### 324-count URL: https://datascientyst.com/count/ Last updated: 2021-11-30T13:46:16.000Z count() ### 320-basic-concepts URL: https://datascientyst.com/basic-concepts-8/ Last updated: 2021-11-30T13:46:15.000Z Basic concepts ### 321-apply URL: https://datascientyst.com/apply/ Last updated: 2021-11-30T13:46:15.000Z apply() ### 322-aggfunc URL: https://datascientyst.com/aggfunc/ Last updated: 2021-11-30T13:46:15.000Z aggfunc ### 313-regex URL: https://datascientyst.com/regex/ Last updated: 2021-11-30T13:46:14.000Z Regex ### 314-search URL: https://datascientyst.com/search/ Last updated: 2021-11-30T13:46:14.000Z Search ### 319-exercise URL: https://datascientyst.com/exercise-7/ Last updated: 2021-11-30T13:46:14.000Z Exercise ### 310-basic-concepts URL: https://datascientyst.com/basic-concepts-7/ Last updated: 2021-11-30T13:46:13.000Z Basic concepts ### 311-replace URL: https://datascientyst.com/replace/ Last updated: 2021-11-30T13:46:13.000Z replace() ### 312-split URL: https://datascientyst.com/split/ Last updated: 2021-11-30T13:46:13.000Z split() ### 233-reshape URL: https://datascientyst.com/reshape/ Last updated: 2021-11-30T13:46:12.000Z Reshape ### 234-melt URL: https://datascientyst.com/melt/ Last updated: 2021-11-30T13:46:12.000Z melt() ### 239-exercise URL: https://datascientyst.com/exercise-6/ Last updated: 2021-11-30T13:46:12.000Z Exercise ### 230-basic-concepts URL: https://datascientyst.com/basic-concepts-6/ Last updated: 2021-11-30T13:46:11.000Z Basic concepts ### 231-groupby URL: https://datascientyst.com/groupby/ Last updated: 2021-11-30T13:46:11.000Z groupby() ### 229-exercise URL: https://datascientyst.com/exercise-5/ Last updated: 2021-11-30T13:46:10.000Z Exercise ### 225-query URL: https://datascientyst.com/query/ Last updated: 2023-02-18T09:01:23.000Z query() ### 226-get URL: https://datascientyst.com/get/ Last updated: 2023-03-04T08:07:53.000Z Get ### 222-find URL: https://datascientyst.com/find/ Last updated: 2021-11-30T13:46:09.000Z Find ### 223-filter URL: https://datascientyst.com/filter/ Last updated: 2021-11-30T13:46:09.000Z Filter ### 219-exercise URL: https://datascientyst.com/exercise-4/ Last updated: 2021-11-30T13:46:08.000Z Exercise ### 220-basic-concepts URL: https://datascientyst.com/basic-concepts-5/ Last updated: 2021-11-30T13:46:08.000Z Basic concepts ### 221-iloc URL: https://datascientyst.com/iloc/ Last updated: 2023-02-18T09:04:55.000Z iloc() ### 213-to-dict URL: https://datascientyst.com/to-dict/ Last updated: 2021-11-30T13:46:07.000Z to\_dict() ### 210-basic-concepts URL: https://datascientyst.com/basic-concepts-4/ Last updated: 2021-11-30T13:46:06.000Z Basic concepts ### 211-to-csv URL: https://datascientyst.com/to-csv/ Last updated: 2021-11-30T13:46:06.000Z to\_csv() ### 134-kaggle URL: https://datascientyst.com/kaggle/ Last updated: 2021-11-30T13:46:05.000Z Kaggle ### 139-exercise URL: https://datascientyst.com/exercise-3/ Last updated: 2021-11-30T13:46:05.000Z Exercise ### 130-basic-concepts URL: https://datascientyst.com/basic-concepts-3/ Last updated: 2021-11-30T13:46:04.000Z Basic concepts ### 131-read-csv URL: https://datascientyst.com/read-csv/ Last updated: 2021-11-30T13:46:04.000Z read\_csv() ### 132-read-excel URL: https://datascientyst.com/read-excel/ Last updated: 2021-11-30T13:46:04.000Z read\_excel() ### 124-multiindex URL: https://datascientyst.com/multiindex/ Last updated: 2021-11-30T13:46:03.000Z MultiIndex ### 129-exercise URL: https://datascientyst.com/exercise-2/ Last updated: 2021-11-30T13:46:03.000Z Exercise ### 123-index URL: https://datascientyst.com/index/ Last updated: 2022-10-13T20:36:46.000Z Index ### 120-basic-concepts URL: https://datascientyst.com/basic-concepts-2/ Last updated: 2021-11-30T13:46:02.000Z Basic concepts ### 121-row URL: https://datascientyst.com/row/ Last updated: 2021-11-30T13:46:02.000Z Row ### 122-column URL: https://datascientyst.com/column/ Last updated: 2021-11-30T13:46:02.000Z Column ### 116-data-types URL: https://datascientyst.com/data-types/ Last updated: 2021-11-30T13:46:01.000Z Data Types ### 119-exercise URL: https://datascientyst.com/exercise/ Last updated: 2021-11-30T13:46:01.000Z Exercise ### 113-series URL: https://datascientyst.com/series/ Last updated: 2021-11-30T13:46:00.000Z Series ### 114-dataframe URL: https://datascientyst.com/dataframe/ Last updated: 2021-11-30T13:46:00.000Z DataFrame ### 115-create URL: https://datascientyst.com/create/ Last updated: 2021-11-30T13:46:00.000Z Create ### 110-basic-concepts URL: https://datascientyst.com/basic-concepts/ Last updated: 2021-11-30T13:45:59.000Z Basic concepts ### 112-installations URL: https://datascientyst.com/installations/ Last updated: 2021-11-30T13:45:59.000Z Installations ### 41-data URL: https://datascientyst.com/data/ Last updated: 2021-11-30T14:56:53.000Z Data ### 42-visualization URL: https://datascientyst.com/visualization/ Last updated: 2021-11-30T14:57:14.000Z Visualization ### 43-challenge URL: https://datascientyst.com/challenge/ Last updated: 2022-09-22T14:20:43.000Z Projects & Challenges ### 31-string-operation URL: https://datascientyst.com/string-operation/ Last updated: 2021-11-30T14:56:01.000Z String operation ### 32-special-operation URL: https://datascientyst.com/special-operation/ Last updated: 2021-11-30T14:56:17.000Z Special operation ### 33-merge-concat URL: https://datascientyst.com/merge-concat/ Last updated: 2023-02-15T07:15:00.000Z Merge & Concat & Pivot ### 21-export URL: https://datascientyst.com/export/ Last updated: 2021-11-30T14:54:58.000Z Export ### 22-access-data URL: https://datascientyst.com/access-data/ Last updated: 2021-11-30T14:55:17.000Z Access Data ### 23-modify-dataframe URL: https://datascientyst.com/modify-dataframe/ Last updated: 2021-11-30T14:55:44.000Z Modify DataFrame ### 11-getting-started URL: https://datascientyst.com/getting-started/ Last updated: 2021-11-30T14:49:39.000Z Getting started ### 12-dataframe-attributes URL: https://datascientyst.com/dataframe-attributes/ Last updated: 2021-11-30T14:54:14.000Z DataFrame Attributes ### 13-import URL: https://datascientyst.com/import/ Last updated: 2021-11-30T14:54:38.000Z Import ### Data Science Challenge 1: Data Styling URL: https://datascientyst.com/data-science-challenge-data-styling/ Last updated: 2021-11-16T15:25:02.000Z Table and Data visualization is a very important part of Data Science. In this challenge we will focus on highlighting data in DataFrame. Challenges have 3 sections depending on your level. Suppose you have DataFrame like: ```python import pandas as pd import numpy as np np.random.seed(24) df = pd.DataFrame({'A': np.linspace(1, 10, 10)}) df = pd.concat([df, pd.DataFrame(np.random.randn(10, 4), columns=list('BCDE'))], axis=1) df.iloc[3, 3] = np.nan df.iloc[0, 2] = np.nan ``` as: | A | B | C | D | E | | --- | ---------- | ---------- | ---------- | ---------- | | 1.0 | 1.329212 | NaN | \-0.316280 | \-0.990810 | | 2.0 | \-1.070816 | \-1.438713 | 0.564417 | 0.295722 | | 3.0 | \-1.626404 | 0.219565 | 0.678805 | 1.889273 | | 4.0 | 0.961538 | 0.104011 | NaN | 0.850229 | | 5.0 | 1.453425 | 1.057737 | 0.165562 | 0.515018 | There are 3 different challenges depending on difficulty level: - Beginner - B1 - Aadvanced - A1 - Master - M1 ## Highlight negative values Can you highlight negative values in red? (as shown below) | | A | B | C | D | E | | - | --------- | ---------- | ---------- | ---------- | ---------- | | 0 | 1.000000 | 1.329212 | nan | \-0.316280 | \-0.990810 | | 1 | 2.000000 | \-1.070816 | \-1.438713 | 0.564417 | 0.295722 | | 2 | 3.000000 | \-1.626404 | 0.219565 | 0.678805 | 1.889273 | | 3 | 4.000000 | 0.961538 | 0.104011 | nan | 0.850229 | | 4 | 5.000000 | 1.453425 | 1.057737 | 0.165562 | 0.515018 | | 5 | 6.000000 | \-1.336936 | 0.562861 | 1.392855 | \-0.063328 | | 6 | 7.000000 | 0.121668 | 1.207603 | \-0.002040 | 1.627796 | | 7 | 8.000000 | 0.354493 | 1.037528 | \-0.385684 | 0.519818 | | 8 | 9.000000 | 1.686583 | \-1.325963 | 1.428984 | \-2.089354 | | 9 | 10.000000 | \-0.129820 | 0.631523 | \-0.586538 | 0.290720 | ## Highlight negative values Can you highlight columns based on the column name? Sample color map: `{'A':'cyan', 'B':'lightblue', 'C':'lightyellow', 'D':'salmon', 'E':'lightgreen'}` | | A | B | C | D | E | | - | --------- | ---------- | ---------- | ---------- | ---------- | | 0 | 1.000000 | 1.329212 | nan | \-0.316280 | \-0.990810 | | 1 | 2.000000 | \-1.070816 | \-1.438713 | 0.564417 | 0.295722 | | 2 | 3.000000 | \-1.626404 | 0.219565 | 0.678805 | 1.889273 | | 3 | 4.000000 | 0.961538 | 0.104011 | nan | 0.850229 | | 4 | 5.000000 | 1.453425 | 1.057737 | 0.165562 | 0.515018 | | 5 | 6.000000 | \-1.336936 | 0.562861 | 1.392855 | \-0.063328 | | 6 | 7.000000 | 0.121668 | 1.207603 | \-0.002040 | 1.627796 | | 7 | 8.000000 | 0.354493 | 1.037528 | \-0.385684 | 0.519818 | | 8 | 9.000000 | 1.686583 | \-1.325963 | 1.428984 | \-2.089354 | | 9 | 10.000000 | \-0.129820 | 0.631523 | \-0.586538 | 0.290720 | ## New column - sum of A & B. Showing split based on the percent Can you add a new column which is sum of both A and B? Then style it by showing split based on the percent of which value from column A or B as shown below: | | A | B | B/A % | total | | - | -- | - | --------- | ----- | | 0 | 12 | 3 | 25.000000 | 15 | | 1 | 6 | 4 | 66.666667 | 10 | | 2 | 8 | 1 | 12.500000 | 9 | | 3 | 15 | 7 | 46.666667 | 22 | ### How to Convert DateTime to Day of Week(name and number) in Pandas URL: https://datascientyst.com/convert-datetime-day-of-week-name-number-in-pandas/ Last updated: 2021-11-11T12:38:39.000Z In this quick tutorial, we'll cover how we **convert datetime to day of week or day name in Pandas.** There are two methods which can get day name of day of the week in Pandas: **(1) Get the name of the day from the week** ```python df['date'].dt.day_name() ``` result: ``` 0 Wednesday 1 Thursday ``` **(2) Get the number of the day from the week** ```python df['date'].dt.day_of_week ``` result: ``` 0 2 1 3 ``` Lets see a simple example which shows both ways. ## Step 1: Create sample DataFrame with date To start let's create a basic DataFrame with date column: ```python import pandas as pd data = {'productivity': [80, 20, 60, 30, 50, 55, 95], 'salary': [3500, 1500, 2000, 1000, 2000, 1500, 4000], 'age': [25, 30, 40, 35, 20, 40, 22], 'duedate': ["2020-10-14", "2020-10-15","2020-10-15", "2020-10-17","2020-10-14","2020-10-14","2020-10-18"], 'person': ['Tim', 'Jim', 'Kim', 'Bim', 'Dim', 'Sim', 'Lim'] } df = pd.DataFrame(data) df ``` result: | productivity | salary | age | duedate | person | | ------------ | ------ | --- | ---------- | ------ | | 80 | 3500 | 25 | 2020-10-14 | Tim | | 20 | 1500 | 30 | 2020-10-15 | Jim | | 60 | 2000 | 40 | 2020-10-15 | Kim | | 30 | 1000 | 35 | 2020-10-17 | Bim | | 50 | 2000 | 20 | 2020-10-14 | Dim | If the date column is stored as a string we need to convert it to datetime by: ```python df['date'] = pd.to_datetime(df['duedate']) ``` You can learn more about datetime conversion in this article: [How to Fix Pandas to\_datetime: Wrong Date and Errors](https://datascientyst.com/how-to-fix-pandas-to%5Fdatetime-wrong-date-and-errors/). ## Step 2: Extract Day of week from DateTime in Pandas To return day names from datetime in Pandas you can use method: `day_name`: ```python df['date'].dt.day_name() ``` Tde result is days of week for each date: ``` 0 Wednesday 1 Thursday 2 Thursday 3 Saturday 4 Wednesday ``` Default language is english. It will return different languages depending on your locale. ## Step 3: Extract Day number of week from DateTime in Pandas If you need to get the day number of the week you can use the following: - `day_of_week` - `dayofweek` \- alias - `weekday` \- alias This method counts days starting from Monday - which is presented by 0. So for our DataFrame above we will have: ```python df['date'].dt.day_of_week ``` day numbers of the week is the output: ``` 0 2 1 3 2 3 3 5 4 2 ``` The final DataFrame will look like: | duedate | person | date | day\_name | day\_number | | ---------- | ------ | ---------- | --------- | ----------- | | 2020-10-14 | Tim | 2020-10-14 | Wednesday | 2 | | 2020-10-15 | Jim | 2020-10-15 | Thursday | 3 | | 2020-10-15 | Kim | 2020-10-15 | Thursday | 3 | | 2020-10-17 | Bim | 2020-10-17 | Saturday | 5 | | 2020-10-14 | Dim | 2020-10-14 | Wednesday | 2 | ![](https://datascientyst.com/content/images/2021/11/convert-datetime-day-of-week-name-number-in-pandas-1.png) ## Step 4: Get day name and number of week If you need to get both of them and use it at the same time. For example to plot bar chart with the natural order of the days - Starting from Monday, Tuesday etc ```python df[['day_number', 'salary', 'day_name']].groupby(['day_number', 'day_name']).mean().sort_index().plot(kind='bar', legend=None) ``` For more information and detailed example you can check this article: [Dates and Bar Plots (per weekday) in Pandas](https://blog.softhints.com/dates-and-bar-plots-in-pandas-day-week/?ref=datascientyst.com) ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/datetime/convert-datetime-day-of-week-name-number-in-pandas.ipynb?ref=datascientyst.com) - [pandas.Series.dt.day\_name](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.dt.day%5Fname.html?ref=datascientyst.com) - [pandas.Series.dt.day\_of\_week](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.dt.day%5Fof%5Fweek.html?ref=datascientyst.com) ### How to Reset Column Names (Index) in Pandas URL: https://datascientyst.com/reset-column-names-index-pandas/ Last updated: 2021-11-05T12:56:46.000Z To **reset column names (column index) in Pandas to numbers from 0 to N** we can use several different approaches: **(1) Range from `df.columns.size`** ```python df.columns = range(df.columns.size) ``` **(2) Transpose to rows and `reset_index` \- the slowest options** ```python df.T.reset_index(drop=True).T ``` **(3) Range from column number - df.shape\[1\]** ```python df.columns = range(df.shape[1]) ``` Which one to use depends on data and the context in which it is used. If you like to change the order of the columns you can check: [How to Change the Order of Columns in Pandas DataFrame](https://datascientyst.com/change-order-columns-pandas-dataframe/) × **Pro Tip** In **Pandas** there are two axes. Rows are considered as indexes. The other one is columns. Information for them is returned from method axes Below you can find simple example and performance comparison: ```python import pandas as pd df = pd.DataFrame({ 'name':['Softhints', 'DataScientyst', 'DataPlotPlus'], 'url':['https://www.softhints.com', 'https://datascientyst.com', 'https://dataplotplus.com'], 'id':['a', 'b', 'c'] }, index=[2, 3, 4]) }) ``` result: | 0 | 1 | 2 | | ------------- | ------------------------- | - | | Softhints | https://www.softhints.com | a | | DataScientyst | https://datascientyst.com | b | | DataPlotPlus | https://dataplotplus.com | c | The columns are: ``` Index(['name', 'url', 'id'], dtype='object') ``` After reset by any of the above we get: ``` RangeIndex(start=0, stop=3, step=1) ``` | 0 | 1 | 2 | | ------------- | ------------------------- | - | | Softhints | https://www.softhints.com | a | | DataScientyst | https://datascientyst.com | b | | DataPlotPlus | https://dataplotplus.com | c | ![](https://datascientyst.com/content/images/2021/11/reset-column-names-index-pandas.png) ## Reset row and column index In order to reset row and column index at the same time you can use Python tuples syntax like: ```python df.index, df.columns = [range(df.index.size), range(df.columns.size)] ``` Prior the reset: | | name | url | id | | - | ------------- | ------------------------- | -- | | 2 | Softhints | https://www.softhints.com | a | | 3 | DataScientyst | https://datascientyst.com | b | | 4 | DataPlotPlus | https://dataplotplus.com | c | After the reset: | | 0 | 1 | 2 | | - | ------------- | ------------------------- | - | | 0 | Softhints | https://www.softhints.com | a | | 1 | DataScientyst | https://datascientyst.com | b | | 2 | DataPlotPlus | https://dataplotplus.com | c | To get axes information from Pandas DataFrame we can use method `axes`: ```python df.axes ``` result before the reset of the indexes: ``` [Int64Index([2, 3, 4], dtype='int64'), RangeIndex(start=0, stop=3, step=1)] ``` ## Performance comparison for resetting column names Lets increase DataFrame rows by: ```python df_perf = pd.concat([df] * 10 ** 4) ``` to: ``` (30000, 3) ``` So the timings are: - `range(df.shape[1])` \- 10.4 µs ± 88.2 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each) - `df.T.reset_index(drop=True).T` \- 449 ms ± 4.38 ms per loop (mean ± std. dev. of 7 runs, 1 loop each) - `df_perf.columns = range(df.columns.size)` \- 10.1 µs ± 77.4 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each) For DataFrame with 30000 columns and shape: (3, 30000): - `range(df.shape[1])` \- 10.4 µs ± 101 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each) - `df.T.reset_index(drop=True).T` \- 444 ms ± 7.38 ms per loop (mean ± std. dev. of 7 runs, 1 loop each) - `df_perf.columns = range(df.columns.size)` \- 10.1 µs ± 126 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each) ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/index/reset-column-names-index-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.reset\_index](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.reset%5Findex.html?ref=datascientyst.com) - [pandas.DataFrame.T](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.T.html?ref=datascientyst.com) - [pandas.DataFrame.axes](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.axes.html?ref=datascientyst.com) ### How to Change the Order of Columns in Pandas DataFrame URL: https://datascientyst.com/change-order-columns-pandas-dataframe/ Last updated: 2021-11-05T08:07:33.000Z Here are two ways to **sort or change the order of columns in Pandas DataFrame.** **(1) Use method `reindex` \- custom sorts** ```python df = df.reindex(sorted(df.columns), axis=1) ``` **(2) Use method `sort_index` \- sort with duplicate column names** ```python df = df.sort_index(axis=1) ``` What is **the difference between if need to change order of columns in DataFrame : `reindex` and `sort_index`**. The `sort_index` is a bit faster (depends on data and column number) and can be used with duplicate names. `reindex` is suitable if you need to apply custom order or sorting. Both of them work for the two axis - rows and columns, Suppose we have data like: | Region | 1500 | 1600 | 1700 | 1750 | 1800 | 1850 | 1900 | | -------------------------- | ---- | ---- | ---- | ---- | ---- | ---- | ---- | | World | 585 | 660 | 710 | 791 | 978 | 1262 | 1650 | | Africa | 86 | 114 | 106 | 106 | 107 | 111 | 133 | | Asia | 282 | 350 | 411 | 502 | 635 | 809 | 947 | | Europe | 168 | 170 | 178 | 190 | 203 | 276 | 408 | | Latin America \[Note 1\] ​ | 40 | 20 | 10 | 16 | 24 | 38 | 74 | Where the full list of columns is: ``` Index(['Region', '1500', '1600', '1700', '1750', '1800', '1850', '1900', '1950', '1999', '2008', '2010', '2012', '2050', '2150'], dtype='object') ``` Data is available by: ```python import pandas as pd df = pd.read_csv('https://raw.githubusercontent.com/softhints/Pandas-Tutorials/master/data/population/population.csv') ``` Before the change of the order let's shuffle the columns and get the initial order: ```python import random initial_order = df.columns.to_list() cols = df.columns.to_list() random.shuffle(cols) ``` ## 1: Change order of columns by `reindex` First example will show us how to use method `reindex` in order to sort the columns in alphabetical order: ```python df = df.reindex(sorted(df.columns), axis=1) ``` By `df.columns` we get all column names as they are stored in the DataFrame. We sort them by `sorted` and finally use the method `reindex` on columns. A shorter code to sort the columns by name would be: ```python df = df[sorted(df.columns)] ``` ### Custom sort of columns with `reindex` In order to change the column order in a custom way we can use method `reindex`. As we saw earlier we can get the list of columns and shuffle them: ```python cols = df.columns.to_list() random.shuffle(cols) ``` Now we can apply this order to the DataFrame by: ```python df = df.reindex(df.columns, axis=1) ``` The columns of the updated DataFrame: ``` Index(['1500', '1900', '2000', '1900', '1850', '2000', '2000', '1900', '1700', '1800', '1600', '2000', '2150', '1750', 'Region'], dtype='object') ``` ![](https://datascientyst.com/content/images/2021/11/change-order-columns-pandas-dataframe.png) ## 2: Sort columns by name by method `sort_index` Alternative solution is to use the method `sort_index`. It doesn't support custom order but it's faster in general. To update the column order by `sort_index` use this syntax: ```python df = df.sort_index(axis=1) ``` The official documentation for this method says: > Returns a new DataFrame sorted by label if inplace argument is False, otherwise updates the original DataFrame and returns None. This method has parameter `inplace` \- which is not the case for `reindex`. ## 3: Sort with duplicate column names Finally let's see what will happen if we apply method `reindex` on DataFrame with duplicate column names. To achieve this we are going to update column names manually: ```python df.columns = ['1500', '1600', '1700', '1750', '1800', '1850', '1900', '1900', '1900', '2000', '2000', '2000', '2000', '2150', 'Region'] ``` Method `reindex` is raising error: > ValueError: cannot reindex from a duplicate axis While `sort_index` is working successfully ## 4: Shift columns in Pandas DataFrame Finally let's see **how to shift columns in Pandas DataFrame.** This is possible by getting a list of columns names and updating the list of columns: ```python cols = df.columns.to_list() cols = cols[-2:] + cols[:-2] ``` result: ``` ['2050', '2150', 'Region', '1500', '1600', '1700', '1750', '1800', '1850', '1900', '1950', '1999', '2008', '2010', '2012'] ``` Finally we can update the DataFrame order by: ```python df = df.reindex(cols, axis=1) ``` or by: ```python df = df[cols] ``` The `df.reindex` is the faster than the second solution ## 5: Performance comparison for `reindex` and `sort_index` Finally lets check the performance for a pretty small DataFrame - (7, 15) between: - `reindex` \- 254 µs ± 1.84 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) - `sort_index` \- 181 µs ± 7.34 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) The same comparison for (700000, 15): - `reindex` \- 24.1 ms ± 740 µs per loop (mean ± std. dev. of 7 runs, 10 loops each) - `sort_index` \- 22.7 ms ± 430 µs per loop (mean ± std. dev. of 7 runs, 10 loops each) For (7, 1500): - `reindex` \- 383 µs ± 4.93 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) - `sort_index` \- 826 µs ± 11.4 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/column/change-order-columns-pandas-dataframe.ipynb?ref=datascientyst.com) - [pandas.DataFrame.reindex](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.reindex.html?ref=datascientyst.com) - [pandas.DataFrame.sort\_index](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.sort%5Findex.html?ref=datascientyst.com) - [Python sorted](https://docs.python.org/3/howto/sorting.html?ref=datascientyst.com) ### How to Search for String in the Whole DataFrame in Pandas URL: https://datascientyst.com/search-for-string-whole-dataframe-pandas/ Last updated: 2021-11-04T09:37:36.000Z To search for a string in all columns of a Pandas DataFrame we can use two different ways: **(1) Lambda and str.contains** ```python df.apply(lambda row: row.astype(str).str.contains('data').any(), axis=1) ``` **(2) np.column\_stack + str.contains** ```python import numpy as np mask = np.column_stack([df[col].astype(str).str.contains("data", na=False) for col in df]) df.loc[mask.any(axis=1)] ``` Let's check two examples on how to use the above techniques in practice. To start with DataFrame like: ```python from IPython.display import HTML import pandas as pd df = pd.DataFrame({ 'id':[1,2,3,4], 'name':['Softhints\nLinux', 'dataplotplus', 'DataScientyst\nPandas', 'test'], 'url':['https://www.softhints.com', 'https://dataplotplus.com/', 'https://datascientyst.com', 'test\data'] }) ``` which has this data: | id | name | url | | -- | ---------------------- | ------------------------- | | 1 | Softhints\\nLinux | https://www.softhints.com | | 2 | dataplotplus | https://dataplotplus.com/ | | 3 | DataScientyst\\nPandas | https://datascientyst.com | | 4 | test | test\\data | ## Search whole DataFrame with lambda and str.contains Searching with lambda and str.contains is straightforward: ```python df.apply(lambda row: row.astype(str).str.contains('data').any(), axis=1) ``` The lambda will iterate over all rows. Then we will convert the values to string - in order to avoid errors. The converted data will be searched for a string pattern - in this case `data`. Without method any we will get the search result for each value: | id | name | url | | ----- | ----- | ----- | | False | False | False | | False | True | True | | False | False | True | | False | False | True | So method `any` will return True if there is at least one True value per row. So the final output is: 0 False 1 True 2 True 3 True dtype: bool ![](https://datascientyst.com/content/images/2021/11/search-for-string-whole-dataframe-pandas.png) ## Search whole DataFrame with numpy and str.contains As an alternative solution you can use the Numpy method - `column_stack` to find all values in all columns. This solution is faster than the previous one. ```python import numpy as np mask = np.column_stack([df[col].astype(str).str.contains("data", na=False) for col in df]) df.loc[mask.any(axis=1)] ``` **How does it work?** So the code: ```python [df[col].astype(str).str.contains("data", na=False) for col in df] ``` will iterate over all columns and then will convert values to string. Then we perform a search for a given value. The result would be: ``` [0 False 1 False 2 False 3 False Name: id, dtype: bool, 0 False 1 True 2 False 3 False Name: name, dtype: bool, 0 False 1 True 2 True 3 True Name: url, dtype: bool] ``` Method `np.column_stack` will convert the above result into `array`: ``` array([[False, False, False], [False, True, True], [False, False, True], [False, False, True]]) ``` Again we are going to use method `any` in order to return a single True or False value per row. The method `loc` will return only the rows which contain the searched value. ```python df.loc[mask.any(axis=1)] ``` ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/find/search-for-string-whole-dataframe-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.apply](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.apply.html?highlight=apply&ref=datascientyst.com#pandas.DataFrame.apply) - [numpy.column\_stack](https://numpy.org/doc/stable/reference/generated/numpy.column%5Fstack.html?ref=datascientyst.com) - [pandas.DataFrame.any](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.any.html?highlight=any&ref=datascientyst.com#pandas.DataFrame.any) - [pandas.Series.any](https://pandas.pydata.org/docs/reference/api/pandas.Series.any.html?highlight=any&ref=datascientyst.com#pandas.Series.any) ### How to Find All Rows With Newline Inside a Column or DataFrame in Pandas URL: https://datascientyst.com/find-all-rows-with-newline-inside-column-dataframe-in-pandas/ Last updated: 2021-11-04T09:08:30.000Z Here is the way **to search for newlines in a column or DataFrame in Pandas**: **(1) Find for newlines in a single column** ```python df[df['name'].str.contains('\n',regex=False)] ``` **(2) Search for newlines in the whole DataFrame** ```python df[df.apply(lambda row: row.astype(str).str.contains('\n').any(), axis=1)] ``` Let say that we have the next DataFrame: | id | name | url | | -- | ---------------------- | ------------------------- | | 1 | Softhints\\nLinux | https://www.softhints.com | | 2 | dataplotplus | https://dataplotplus.com/ | | 3 | DataScientyst\\nPandas | https://datascientyst.com | | 4 | test | test\\ntest | ## Search for newlines in single column If you need to return only the rows which contains newlines - `\n` in a single column - `name`. Then we can use string method - `contains` of Pandas in combination with `regex=False`: ```python df[df['name'].str.contains('\n',regex=False)] ``` the returned rows are: | id | name | url | | -- | ---------------------- | ------------------------- | | 1 | Softhints\\nLinux | https://www.softhints.com | | 3 | DataScientyst\\nPandas | https://datascientyst.com | ## Search for newlines in multiple columns of DataFrame If you need to **search for newlines and other break symbols in several columns** \- you can use lambda. First we will convert the values to string - in order to avoid errors for non string values. Then we are going to search for `\n`: ```python df[df[['name', 'url']].apply(lambda row: row.astype(str).str.contains('\n').any(), axis=1)] ``` rows which has newlines in those two columns: | id | name | url | | -- | ---------------------- | ------------------------- | | 1 | Softhints\\nLinux | https://www.softhints.com | | 3 | DataScientyst\\nPandas | https://datascientyst.com | | 4 | test | test\\ntest | ## Search for newlines in multiple columns of DataFrame **Searching the whole Pandas DataFrame for breaks and newlines again uses lambda**: ```python df[df.apply(lambda row: row.astype(str).str.contains('\n').any(), axis=1)] ``` result: | id | name | url | | -- | ---------------------- | ------------------------- | | 1 | Softhints\\nLinux | https://www.softhints.com | | 3 | DataScientyst\\nPandas | https://datascientyst.com | | 4 | test | test\\ntest | ![](https://datascientyst.com/content/images/2021/11/find-all-rows-with-newline-inside-column-dataframe-in-pandas.png) ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/commit/2fedaa2f78f63bd0e59b337269d546c5253a5bbe?ref=datascientyst.com) - [pandas.Series.str.contains](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.contains.html?highlight=contains&ref=datascientyst.com#pandas.Series.str.contains) - [pandas.DataFrame.apply](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.apply.html?highlight=apply&ref=datascientyst.com#pandas.DataFrame.apply) ### How to Pretty Print Newlines in Pandas DataFrame URL: https://datascientyst.com/pretty-print-newlines-in-pandas-dataframe/ Last updated: 2023-09-26T14:18:15.000Z In this short post, you'll see how to pretty print newlines inside the values in Pandas DataFrame. Brief example is included for demonstration purposes. You can find the two ways below: **(1) Using df.to\_html().replace** ```python from IPython.display import display, HTML display( HTML( df.to_html().replace("\\n","
"))) ``` **(2) By df.style.set\_properties** ```python display(df.style.set_properties(**{ 'text-align': 'left', 'white-space': 'pre-wrap', }) ``` Suppose we have a DataFrame like: | name | url | url2 | | ------------- | ------------------------- | ----------------------------------------- | | Softhints | https://www.softhints.com | https://www.blog.softhints.com/tag/pandas | | DataScientyst | https://datascientyst.com | https://datascientyst.com/tag/pandas | Let's create new column which is concatenation of all others with separator a new line: ```python df['all'] = df['name'] + '\n' + df['url'] + '\n' + df['url2'] ``` result: | url2 | all | | ----------------------------------------- | --------------------------------------------------------------------------------- | | https://www.blog.softhints.com/tag/pandas | Softhints\\nhttps://www.softhints.com\\nhttps://www.blog.softhints.com/tag/pandas | | https://datascientyst.com/tag/pandas | DataScientyst\\nhttps://datascientyst.com\\nhttps://datascientyst.com/tag/pandas | As you can see default print is showing the string values as a single line - ignoring the newlines. ## df.to\_html().replace If you like to print newlines you need to use the following approach: To print or display new lines inside columns in Pandas DataFrame you can: \* convert the output to HTML - replace all new lines with `
` \- HTML tag for new line - display the result as HTML ```python from IPython.display import display, HTML display( HTML( df.to_html().replace("\\n","
") ) ) ``` now all **newlines symbols are printed or displayed in Jupyter cells correctly**: | | url2 | all | | - | ----------------------------------------- | --------------------------------------------------------------------------- | | 0 | https://www.blog.softhints.com/tag/pandas | Softhintshttps://www.softhints.comhttps://www.blog.softhints.com/tag/pandas | | 1 | https://datascientyst.com/tag/pandas | DataScientysthttps://datascientyst.comhttps://datascientyst.com/tag/pandas | You can define also a method like: ```python from IPython.display import display, HTML def print_newlines(df): return display( HTML( df.to_html().replace("\\n","\") ) ) print_newlines(df) ``` ## df.style.set\_properties Alternative method is by changing Pandas properties: - `white-space` - `text-align` Let say you would like to have newlines and centered values. Then you can use `style.set_properties` with the following values: ```python df.style.set_properties(**{ 'text-align': 'center', 'white-space': 'pre-wrap', }) ``` ![](https://datascientyst.com/content/images/2021/11/pretty-print-newlines-in-pandas-dataframe.png) ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/styling/pretty-print-newlines-in-pandas-dataframe.ipynb?ref=datascientyst.com) - [pandas.DataFrame.to\_html](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to%5Fhtml.html?ref=datascientyst.com) - [pandas.io.formats.style.Styler.set\_properties](https://pandas.pydata.org/docs/reference/api/pandas.io.formats.style.Styler.set%5Fproperties.html?highlight=set%5Fproperties&ref=datascientyst.com) ### How to Replace Regex Groups in Pandas URL: https://datascientyst.com/replace-regex-groups-in-pandas/ Last updated: 2021-11-03T14:32:42.000Z In this short tutorial, we'll look at how to **match and replace regex groups in Pandas.** Here you can find the short answer: ```python df_e['Date'].str.replace(r'(\d{2})/(\d{2})/(\d{4})', r"\3-\2-\1", regex=True) ``` ## How to match and replace regex groups in Pandas Let's check first for the regex pattern `r'(\d{2})/(\d{2})/(\d{4})'`. It has 3 capturing groups:: - 1st capturing group `(\d{2})` - `\d` \- matches a digit (equivalent to \[0-9\]) - {2} matches the previous token exactly 2 times - 2nd capturing group `(\d{2})` - `\d` \- matches a digit (equivalent to \[0-9\]) - {2} matches the previous token exactly 2 times - 4nd capturing group `(\d{4})` - `\d` \- matches a digit (equivalent to \[0-9\]) - {4} matches the previous token exactly 4 times Between the capturing groups there are separators - `/` which will be matched but not captured. Now in the replacement part - `r"\3-\2-\1"` \- the groups are presented by numbers - starting from 1: - 1st group is `\1` - 2nd group is `\2` - 3rd group is `\3` So if the initial value in Pandas column is: ``` 0 01/02/1965 1 01/04/1965 2 01/05/1965 3 01/08/1965 4 01/09/1965 ``` After the replacement we will end with: ``` 0 1965-02-01 1 1965-04-01 2 1965-05-01 3 1965-08-01 4 1965-09-01 ``` Another example is: ```python df_e['Date'].str.replace(r'(\d{2})/(\d{2})/\d{2}(\d{2})', r"\2 \1 '\3", regex=True) ``` result will be: ``` 0 02 01 '65 1 04 01 '65 2 05 01 '65 3 08 01 '65 4 09 01 '65 ``` ![](https://datascientyst.com/content/images/2021/11/replace-regex-groups-in-pandas.png) ## How to match and replace regex groups - string patterns Let's have string data which has a different number of words. Suppose we like to extract only the last word from all rows: ``` 0 Noida 1 Noida 2 Work From Home 3 Work From Home 4 Work From Home ... ``` We can use regex groups to match the last word by: ```python df['location'].str.replace(r'(.*) (.*) (.*)', r"\3", regex=True) ``` result: ``` 0 Noida 1 Noida 2 Home 3 Home 4 Home ``` If you need to work with special characters and complex expressions - you may need to test a lot - in order to avoid problems. Let's have data like: ``` 0 Software Testing 1 Java, SQL, Unix, Oracle, MS SQL Server, Hibern... 2 English Proficiency (Spoken), English Proficie... 3 HTML, CSS, Flask, Python, Django 4 HTML, CSS, JavaScript, ReactJS, Redux ``` For example if you expect: ```python df['skills'].str.replace(r'(.*)(\(.*\))(.*)', r"\1\3", regex=True) ``` The regex above to remove all words surrounded by `(` and `)` \- this will not happen: ``` 0 Software Testing 1 Java, SQL, Unix, Oracle, MS SQL Server, Hibern... 2 English Proficiency (Spoken), English Proficie... 3 HTML, CSS, Flask, Python, Django 4 HTML, CSS, JavaScript, ReactJS, Redux ``` because the regex are greedy by default and will try to match as much as possible. You can check what is captured as group 2: ```python df['skills'].str.replace(r'(.*)(\(.*\))(.*)', r"\2", regex=True) ``` So as you can if there are multiple matches - the regex will try to match everything until the end - or matching the last one possible: ``` 0 Software Testing 1 (Java) 2 (Written) 3 HTML, CSS, Flask, Python, Django 4 HTML, CSS, JavaScript, ReactJS, Redux ``` For non greedy matching in Pandas regex groups you can add `?`. So: - `(.*)` \- greedy match - `(.*)` \- non greedy match ```python df['skills'].str.replace(r'(.*?)(\(.*?\))(.*)', r"\2", regex=True) ``` result: ``` 0 Software Testing 1 (Java) 2 (Spoken) 3 HTML, CSS, Flask, Python, Django 4 HTML, CSS, JavaScript, ReactJS, Redux ``` Note that the third line changed from `(Written)` to `(Spoken)`. ## How to match and replace regex groups - numeric patterns Suppose we have ranged amounts like: ``` 0 8000 /month 1 10000 /month 2 1000-2000 /month 3 1000-2000 /month 4 1000-2000 /month ``` And we need to get only the upper amount from this range and leave the rest as it is. This can be achieved by: ```python df['stipend'].str.replace(r'(\d+)-(\d+)(.*)', r"\2\3", regex=True) ``` result: ``` 0 8000 /month 1 10000 /month 2 2000 /month 3 2000 /month 4 2000 /month ``` ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/regex/replace-regex-groups-in-pandas.ipynb?ref=datascientyst.com) - [pandas.Series.str.replace](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.replace.html?ref=datascientyst.com) - [pandas.Series.replace](https://pandas.pydata.org/docs/reference/api/pandas.Series.replace.html?ref=datascientyst.com) ### How to Drop Bad Lines with read_csv in Pandas URL: https://datascientyst.com/drop-bad-lines-with-read_csv-pandas/ Last updated: 2026-07-22T22:19:53.000Z Here are two approaches to **drop bad lines with `read_csv` in Pandas**, when `on_bad_lines` is the only option, since: - `error_bad_lines` - `warn_bad_lines` were removed in pandas 2.0. **(1) Parameter `on_bad_lines='skip'` — Pandas >= 1.3 (current, recommended)** ```python df = pd.read_csv(csv_file, delimiter=';', on_bad_lines='skip') ``` **(2) `error_bad_lines=False` — Pandas < 1.3 (removed, do not use)** ```python # ❌ No longer works — raises TypeError on any pandas 2.x install df = pd.read_csv(csv_file, delimiter=';', error_bad_lines=False) ``` `error_bad_lines` and `warn_bad_lines` were deprecated back in pandas 1.3 and fully **removed in pandas 2.0.0**. If you still have this in an older script, calling it on a modern pandas version will raise `TypeError: read_csv() got an unexpected keyword argument 'error_bad_lines'`. There's no fallback or compatibility shim — you need to switch to `on_bad_lines`. Suppose we have two files: - Single separator `;` ``` Date;Company A;Company A;Company B;Company B 2021-09-06;1;7.9;2;6 2021-09-07;1;8.5;2;7 2021-09-08;2;8;1;8.1 2021-09-09;2;8;1;"8.3;5.5" ``` - Double separator `;;` ``` Date;;Company A;;Company A;;Company B;;Company B 2021-09-06;;1;;7.9;;2;;6 2021-09-07;;1;;8.5;;2;;7 2021-09-08;;2;;8;;1;;8.1 2021-09-09;;2;;8;;1;;"8.3;;5.5" ``` ## The `on_bad_lines` parameter ``` on_bad_lines{'error', 'warn', 'skip'} or callable, default 'error' ``` Specifies what to do upon encountering a bad line (a line with too many fields). Allowed values are: - `error` — raise an exception when a bad line is encountered (this is still the default). - `warn` — raise a warning when a bad line is encountered and skip that line. - `skip` — skip bad lines without raising or warning when they are encountered. - **callable** — since pandas 1.4, you can also pass a function. It receives the bad line as a list of strings and can return a corrected list of fields (to fix the row instead of dropping it), or `None` to drop it. This only works with the `engine='python'` parser. Example of the callable form, useful when you want to repair rather than discard a malformed row: ```python def fix_row(bad_line): # keep only the first 5 fields, drop the rest return bad_line[:5] df = pd.read_csv(csv_file, delimiter=';', engine='python', on_bad_lines=fix_row) ``` Note that depending on the separator: - single - multiple - regex the `read_csv` behavior can be different. You can check this article for more information: [How to Use Multiple Char Separator in read\_csv in Pandas](https://datascientyst.com/use-multiple-char-separator-read%5Fcsv-pandas/) The reason for this is described in the pandas documentation: > Note that regex delimiters are prone to ignoring quoted data. Regex example: `'\r\t'`. Any separator longer than one character (like our `;;` example) is treated as a regex by the parser, which is why quoted fields containing the delimiter get mis-split — this behavior is unchanged in current pandas. For example, for a single separator `";"` this code will work fine: | Date | Company A | Company A.1 | Company B | Company B.1 | | ---------- | --------- | ----------- | --------- | ----------- | | 2021-09-06 | 1 | 7.9 | 2 | 6 | | 2021-09-07 | 1 | 8.5 | 2 | 7 | | 2021-09-08 | 2 | 8.0 | 1 | 8.1 | | 2021-09-09 | 2 | 8.0 | 1 | 8.3;5.5 | While if you have a file with two separators you will get an error or warning (depending on `on_bad_lines`) for the quoted line: ```python csv_file = '../data/csv/multine_bad_line_multi_sep.csv' df = pd.read_csv(csv_file, delimiter=';;', engine='python', on_bad_lines='warn') ``` > Skipping line 5: Expected 5 fields in line 5, saw 6\. Error could possibly be due to quotes being ignored when a multi-char delimiter is used. In this case only 3 rows will be read from the CSV file: | Date | Company A | Company A.1 | Company B | Company B.1 | | ---------- | --------- | ----------- | --------- | ----------- | | 2021-09-06 | 1 | 7.9 | 2 | 6.0 | | 2021-09-07 | 1 | 8.5 | 2 | 7.0 | | 2021-09-08 | 2 | 8.0 | 1 | 8.1 | ![](https://datascientyst.com/content/images/2021/11/drop-bad-lines-with-read_csv-pandas.png) ## A quick note on the C vs Python engine With a single-character delimiter, pandas defaults to the fast C engine, which supports `on_bad_lines='error'|'warn'|'skip'` natively. With a multi-character or regex delimiter, pandas automatically falls back to `engine='python'`, which is slower but is also the only engine that supports a **callable** passed to `on_bad_lines`. If you need to fix (not just skip) bad lines and you're on a single-character delimiter, you must explicitly pass `engine='python'` yourself, as shown above. ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/read%5Fcsv/drop-bad-lines-with-read%5Fcsv-pandas.ipynb?ref=datascientyst.com) - [pandas.read\_csv](https://pandas.pydata.org/docs/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) ### How to Use Multiple Char Separator in read_csv in Pandas URL: https://datascientyst.com/use-multiple-char-separator-read_csv-pandas/ Last updated: 2021-11-03T12:09:57.000Z Here is the way to use **multiple separators (regex separators) with read\_csv in Pandas:** ```python df = pd.read_csv(csv_file, sep=';;', engine='python') ``` Suppose we have a CSV file with the next data: ``` Date;;Company A;;Company A;;Company B;;Company B 2021-09-06;;1;;7.9;;2;;6 2021-09-07;;1;;8.5;;2;;7 2021-09-08;;2;;8;;1;;8.1 multine_separators ``` As you can see there are multiple separators between the values - `;;`. In order to read this we need to specify that as a parameter - `delimiter=';;',`. × **Pro Tip 1** Parameter **delimiter** is alias for **sep** From the official docs: > sepstr, default ‘,’ > Delimiter to use. If sep is None, the C engine cannot automatically detect the separator, but the Python parsing engine can, meaning the latter will be used and automatically detect the separator by Python’s builtin sniffer tool, csv.Sniffer. In this post we are interested mainly in this part: > In addition, separators longer than 1 character and different from '\\s+' will be interpreted as regular expressions and will also force the use of the Python parsing engine. Note that regex delimiters are prone to ignoring quoted data. Regex example: '\\r\\t'. If you try to read the above file without specifying the engine like: ```python df = pd.read_csv(csv_file, delimiter=';;') ``` you will get warning message like: > /home/vanx/PycharmProjects/datascientyst/venv/lib/python3.8/site-packages/pandas/util/\_decorators.py:311: ParserWarning: Falling back to the 'python' engine because the 'c' engine does not support regex separators (separators > 1 char and different from '\\s+' are interpreted as regex); you can avoid this warning by specifying engine='python'. > return func(\*args, \*\*kwargs) So you need to specify the engine like: ```python df = pd.read_csv(csv_file, delimiter=';;', engine='python') ``` What does it mean: > Note that regex delimiters are prone to ignoring quoted data. Regex example: '\\r\\t'. Let's add the following line to the CSV file: ``` 2021-09-09;;2;;8;;1;;"8.3;;5.5" ``` If we try to read this file again we will get an error: > ParserError: Expected 5 fields in line 5, saw 6\. Error could possibly be due to quotes being ignored when a multi-char delimiter is used. You can skip lines which cause errors like the one above by using parameter: `error_bad_lines=False` or `on_bad_lines` for Pandas > 1.3. ```python df = pd.read_csv(csv_file, delimiter=';;', engine='python', error_bad_lines=False) ``` Finally in order to use **regex separator in Pandas:** you can write: ```python df = pd.read_csv(csv_file, sep=r';+', engine='python') ``` ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/read%5Fcsv/use-multiple-char-separator-read%5Fcsv-pandas.ipynb?ref=datascientyst.com) - [pandas.read\_csv](https://pandas.pydata.org/pandas-docs/dev/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) ### Combine Multiple columns into a single one in Pandas URL: https://datascientyst.com/combine-multiple-columns-into-single-one-in-pandas/ Last updated: 2025-12-06T10:12:43.000Z In this short guide, you'll see how to combine multiple columns into a single one in Pandas. Here you can find the short answer: **(1) String concatenation** ```python df['Magnitude Type'] + ', ' + df['Type'] ``` **(2) Using methods `agg` and `join`** ```python df[['Date', 'Time']].T.agg(','.join) ``` **(3) Using lambda and join** ```python df[['Date', 'Time']].agg(lambda x: ','.join(x.values), axis=1).T ``` So let's see several useful examples on how to combine several columns into one with Pandas. Suppose you have data like: | Date | Time | Depth | Magnitude Type | Type | Magnitude | | ---------- | -------- | ----- | -------------- | ---------- | --------- | | 01/02/1965 | 13:44:18 | 131.6 | MW | Earthquake | 6.0 | | 01/04/1965 | 11:29:49 | 80.0 | MW | Earthquake | 5.8 | | 01/05/1965 | 18:05:58 | 20.0 | MW | Earthquake | 6.2 | | 01/08/1965 | 18:49:43 | 15.0 | MW | Earthquake | 5.8 | | 01/09/1965 | 13:32:50 | 15.0 | MW | Earthquake | 5.8 | ## 1: Combine multiple columns using string concatenation Let's start with most simple example - to combine two string columns into a single one separated by a comma: ```python df['Magnitude Type'] + ', ' + df['Type'] ``` result will be: ``` 0 MW, Earthquake 1 MW, Earthquake 2 MW, Earthquake 3 MW, Earthquake 4 MW, Earthquake ``` What if one of the columns is not a string? Then you will get error like: > TypeError: can only concatenate str (not "float") to str To avoid this error you can convert the column by using method `.astype(str)`: ```python df['Magnitude Type'] + ', ' + df['Magnitude'].astype(str) ``` result: ``` 0 MW, 6.0 1 MW, 5.8 2 MW, 6.2 3 MW, 5.8 4 MW, 5.8 ``` ## 2: Combine date and time columns into DateTime column What if you have separate columns for the date and the time. You can concatenate them into a single one by using string concatenation and conversion to datetime: ```python pd.to_datetime(df['Date'] + ' ' + df['Time'], errors='ignore') ``` In case of missing or incorrect data we will need to add parameter: `errors='ignore'` in order to avoid error: > ParserError: Unknown string format: 1975-02-23T02:58:41.000Z 1975-02-23T02:58:41.000Z ![](https://datascientyst.com/content/images/2021/11/combine-multiple-columns-into-single-one-in-pandas.png) ## 3: Combine multiple columns with agg and join Another option to concatenate multiple columns is by using two Pandas methods: - `agg` - `join` ```python df[['Date', 'Time']].T.agg(','.join) ``` result: ``` 0 01/02/1965,13:44:18 1 01/04/1965,11:29:49 2 01/05/1965,18:05:58 3 01/08/1965,18:49:43 ``` This one might be a bit slower than the first one. ## 4: Combine multiple columns with lambda and join You can use lambda expressions in order to concatenate multiple columns. The advantages of this method are several: - you can have condition on your input - like filter - output can be customised - better control on dtypes To combine columns date and time we can do: ```python df[['Date', 'Time']].agg(lambda x: ','.join(x.values), axis=1).T ``` In the next section you can find how we can use this option in order to combine columns with the same name. ## 5: Combine columns which have the same name Finally let's combine all columns which have exactly the same name in a Pandas DataFrame. First let's create duplicate columns by: ```python df.columns = ['Date', 'Date', 'Depth', 'Magnitude Type', 'Type', 'Magnitude'] df ``` A general solution which **concatenates columns with duplicate names can be:** ```python df.groupby(df.columns, axis=1).agg(lambda x: x.apply(lambda y: ','.join([str(l) for l in y if str(l) != "nan"]), axis=1)) ``` This will result into: | Date | Depth | Magnitude | Magnitude Type | Type | | ------------------- | ----- | --------- | -------------- | ---------- | | 01/02/1965,13:44:18 | 131.6 | 6.0 | MW | Earthquake | | 01/04/1965,11:29:49 | 80.0 | 5.8 | MW | Earthquake | | 01/05/1965,18:05:58 | 20.0 | 6.2 | MW | Earthquake | | 01/08/1965,18:49:43 | 15.0 | 5.8 | MW | Earthquake | | 01/09/1965,13:32:50 | 15.0 | 5.8 | MW | Earthquake | How does it work? First is grouping the columns which share the same name: ```python for i in df.groupby(df.columns, axis=1): print(i) ``` result: ``` ('Date', Date Date 0 01/02/1965 13:44:18 1 01/04/1965 11:29:49 2 01/05/1965 18:05:58 3 01/08/1965 18:49:43 4 01/09/1965 13:32:50 ... ... ... 23407 12/28/2016 08:22:12 23408 12/28/2016 09:13:47 23409 12/28/2016 12:38:51 23410 12/29/2016 22:30:19 23411 12/30/2016 20:08:28 [23412 rows x 2 columns]) ('Depth', Depth 0 131.60 1 80.00 2 20.00 3 15.00 4 15.00 ... ... 23407 12.30 23408 8.80 ``` Then it's combining their values: ```python df.groupby(df.columns, axis=1).apply(lambda x: x.values) ``` result: ``` Date [[01/02/1965, 13:44:18], [01/04/1965, 11:29:49... Depth [[131.6], [80.0], [20.0], [15.0], [15.0], [35.... Magnitude [[6.0], [5.8], [6.2], [5.8], [5.8], [6.7], [5.... Magnitude Type [[MW], [MW], [MW], [MW], [MW], [MW], [MW], [MW... Type [[Earthquake], [Earthquake], [Earthquake], [Ea... ``` Finally there is prevention of errors in case of bad values like NaN, missing values, None, different formats etc. ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/column/combine-multiple-columns-into-single-one-in-pandas.ipynb?ref=datascientyst.com) - [Working with text data](https://pandas.pydata.org/docs/user%5Fguide/text.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.agg](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.agg.html?highlight=agg&ref=datascientyst.com#pandas.core.groupby.GroupBy.agg) - [pandas.core.groupby.GroupBy.apply](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.apply.html?highlight=apply&ref=datascientyst.com#pandas.core.groupby.GroupBy.apply) ### Dump (unique) values to CSV / to_csv in Pandas URL: https://datascientyst.com/dump-unique-values-to-csv-to_csv-pandas/ Last updated: 2021-11-03T07:10:45.000Z This post will show you how **to use to\_csv in Pandas** and avoid errors like: > AttributeError: 'numpy.ndarray' object has no attribute 'to\_csv' For example if you like to write unique values from Pandas to a CSV file by method `to_csv` from this data: | Date | Time | Latitude | Longitude | Depth | Magnitude Type | | ---------- | -------- | -------- | --------- | ----- | -------------- | | 01/02/1965 | 13:44:18 | 19.246 | 145.616 | 131.6 | MW | | 01/04/1965 | 11:29:49 | 1.863 | 127.352 | 80.0 | MW | | 01/05/1965 | 18:05:58 | \-20.579 | \-173.972 | 20.0 | MW | | 01/08/1965 | 18:49:43 | \-59.076 | \-23.557 | 15.0 | MW | | 01/09/1965 | 13:32:50 | 11.938 | 126.427 | 15.0 | MW | by: ```python df['Magnitude Type'].unique().to_csv() ``` This will result into: > AttributeError: 'numpy.ndarray' object has no attribute 'to\_csv' The reason for the error is that method `unique` returns an array and not pd.Series or pd.DataFrame object: ``` array(['MW', 'ML', 'MH', 'MS', 'MB', 'MWC', 'MD', nan, 'MWB', 'MWW', 'MWR'], dtype=object) ``` For Pandas you have two different option depending on your context. ### Series Convert the result of method `unique` to a Pandas Series by `pd.Series`: ```python pd.Series(df['Magnitude Type'].unique()).to_csv('data.csv') ``` ### DataFrame If the `pd.Series` it's not suitable or doesn't work for your case then you can use `pd.DataFrame`: ```python pd.DataFrame(df['Magnitude Type'].unique()).to_csv('data.csv') ``` ### Skip headers and index Finally if you like to exclude the headers and the index from the output CSV file you can use arguments: `header=None, index=None` ```python pd.DataFrame(df['Magnitude Type'].unique()).to_csv('data.csv', header=None, index=None) ``` ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/to%5Fcsv/dump-unique-values-to-csv-to%5Fcsv-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.to\_csv](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to%5Fcsv.html?highlight=to%5Fcsv&ref=datascientyst.com#pandas.DataFrame.to%5Fcsv) - [pandas.Series.to\_csv](https://pandas.pydata.org/docs/reference/api/pandas.Series.to%5Fcsv.html?highlight=to%5Fcsv&ref=datascientyst.com#pandas.Series.to%5Fcsv) ### SQL Equivalent count(distinct) in Pandas - nunique URL: https://datascientyst.com/sql-equivalent-count-distinct-pandas-nunique/ Last updated: 2022-06-20T14:12:26.000Z In this post, we will see how to **count distinct values in Pandas**. The SQL's Equivalent of **`count(distinct)` is method `nunique()`. It will count distinct values per group in Pandas DataFrame:** **(1) Pandas method `nunique()` and \`groupby()** ```python df.groupby('Magnitude Type')['Date'].nunique() ``` **(2) Pandas count distinct - `nunique()` \+ `agg()`** ```python df.groupby(['Magnitude Type'])['Date'].agg(['count', 'nunique']) ``` The above solution will group by column 'Magnitude Type' and will **count all unique values for column** 'Date'. If you need to learn more about Pandas and SQL you can visit this cheat sheet: [Pandas vs SQL Cheat Sheet](https://datascientyst.com/pandas-vs-sql-cheat-sheet/). Let's cover the above in more detail. Suppose we have DataFrame like: | Date | Latitude | Longitude | Depth | Magnitude Type | | ---------- | -------- | --------- | ----- | -------------- | | 01/02/1965 | 19.246 | 145.616 | 131.6 | MW | | 01/04/1965 | 1.863 | 127.352 | 80.0 | MW | | 01/05/1965 | \-20.579 | \-173.972 | 20.0 | MW | | 01/08/1965 | \-59.076 | \-23.557 | 15.0 | MW | | 01/09/1965 | 11.938 | 126.427 | 15.0 | MW | ## Step 1: `nunique()` and `groupby()` \- SQL's equivalent to `count(distinct)` in Pandas If we like to group by a column and then **get unique values for another column in Pandas** you can use the method: `.nunique()`. The syntax is simple: ```python df.groupby('Magnitude Type')['Date'].nunique() ``` the result is similar to the SQL query: ```sql SELECT count(distinct Date) FROM earthquakes GROUP BY `Magnitude Type`; ``` ``` Magnitude Type MB 2474 MD 5 MH 4 ML 60 MS 1316 MW 4665 MWB 2036 MWC 3458 MWR 18 MWW 1256 Name: Date, dtype: int64 ``` × **Pro Tip 1** Method **value\_counts** will return the total values and not the unique ones. ```python df['Magnitude Type'].value_counts(dropna=False) ``` result: ``` MW 7722 MWC 5669 MB 3761 MWB 2458 MWW 1983 MS 1702 ML 77 MWR 26 MD 6 MH 5 NaN 3 Name: Magnitude Type, dtype: int64 ``` × **Pro Tip 2** Be careful for **NaN** values. They can make difference in the numbers. ![](https://datascientyst.com/content/images/2021/11/sql-equivalent-count-distinct-pandas-nunique.png) ## Step 2: `nunique()` and `agg()` in Pandas If we like to **get all unique values in each group plus information** like: - `min` - `max` - `count` - `mean` You can modify the upper query a bit: ```python df.groupby(['Magnitude Type'])['Date'].agg(['count', 'nunique']) ``` This will result into: | | count | nunique | min | | -------------- | ----- | ------- | ---------- | | Magnitude Type | | | | | MB | 3761 | 2474 | 01/01/1975 | | MD | 6 | 5 | 02/28/2001 | | MH | 5 | 4 | 02/09/1971 | | ML | 77 | 60 | 01/03/1976 | | MS | 1702 | 1316 | 01/01/1973 | | MW | 7722 | 4665 | 01/01/1967 | | MWB | 2458 | 2036 | 01/01/1995 | | MWC | 5669 | 3458 | 01/01/1997 | | MWR | 26 | 18 | 02/13/2012 | | MWW | 1983 | 1256 | 01/01/2011 | As you can see for `Magnitude Type` \= MH there are 4 unique Dates but 5 records: | | Date | Latitude | Longitude | Depth | Magnitude Type | | ----- | ---------- | --------- | ------------ | ------ | -------------- | | 1848 | 02/09/1971 | 34.416000 | \-118.370000 | 6.000 | MH | | 1849 | 02/09/1971 | 34.416000 | \-118.370000 | 6.000 | MH | | 9664 | 10/18/1989 | 37.036167 | \-121.879833 | 17.214 | MH | | 10584 | 08/17/1991 | 41.679000 | \-125.856000 | 1.303 | MH | | 10869 | 04/25/1992 | 40.335333 | \-124.228667 | 9.856 | MH | Because there are 2 records for date: 02/09/1971. ## Step 3: Pandas count distinct multiple columns To **count number of unique values for multiple columns in Pandas** there are two options: - Combine `agg()` \+ `nunique()` - Pandas method `crosstab()` ### Count distinct with agg() + nunique() First one is using the same approach as above: ```python df.groupby('Magnitude Type').agg({'Date': ['nunique', 'count'], 'Depth': ['nunique', 'count']}) ``` result: | | Date | Depth | | | | -------------- | ------- | ----- | ------- | ----- | | | nunique | count | nunique | count | | Magnitude Type | | | | | | MB | 2474 | 3761 | 911 | 3761 | | MD | 5 | 6 | 6 | 6 | | MH | 4 | 5 | 4 | 5 | | ML | 60 | 77 | 67 | 77 | | MS | 1316 | 1702 | 206 | 1702 | ### Count distinct in Pandas with crosstab() Usually Pandas provide multiple ways of achieving something. In this example we are going to use the method `crosstab`. It might be slower than the other but you may have additional useful information for more insights. So first let's do a simple demo of `crosstab`: ```python pd.crosstab(df['Magnitude Type'], df['Date']) ``` This will give a table which is showing information for each Date and Magnitude Type: | Date | 01/01/1967 | 01/01/1969 | 01/01/1970 | 01/01/1971 | 01/01/1972 | | -------------- | ---------- | ---------- | ---------- | ---------- | ---------- | | Magnitude Type | | | | | | | MW | 2 | 1 | 1 | 1 | 1 | | MWB | 0 | 0 | 0 | 0 | 0 | | MWC | 0 | 0 | 0 | 0 | 0 | | MWR | 0 | 0 | 0 | 0 | 0 | | MWW | 0 | 0 | 0 | 0 | 0 | And then we can summarize the information into unique values by using `.ne(0).sum(1)` : ```python pd.crosstab(df['Magnitude Type'], df['Date']).ne(0).sum(1) ``` result: ``` Magnitude Type MB 2474 MD 5 MH 4 ML 60 MS 1316 MW 4665 MWB 2036 MWC 3458 MWR 18 MWW 1256 dtype: int64 ``` × **How does it work?** Method **ne(0)** will convert anything different from 0 to True and 0 to False. Then we are going to sum across rows. ## Step 4: Count distinct with condition in Pandas Finally we have an option of using a lambda in order to **count the unique values with condition**. This will be useful if you like to apply conditions on the count - for example excluding some records. ```python df.groupby('Magnitude Type')['Date'].apply(lambda x: x.unique().shape[0]) ``` For example example excluding groups smaller than 1000: ```python df.groupby('Magnitude Type')['Date'].apply(lambda x: x.unique().shape[0] if x.unique().shape[0] > 1000 else None).dropna() ``` ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/count/sql-equivalent-count-distinct-pandas-nunique.ipynb?ref=datascientyst.com) - [pandas.DataFrame.nunique](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.nunique.html?ref=datascientyst.com) - [pandas.DataFrame.ne](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.ne.html?ref=datascientyst.com) - [pandas.DataFrame.agg](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.agg.html?ref=datascientyst.com) ### How to Replace Text in a Pandas DataFrame Or Column URL: https://datascientyst.com/replace-text-pandas-dataframe-column/ Last updated: 2021-11-02T11:53:35.000Z **Replace text** is one of the most popular operation in **Pandas DataFrames and columns.** In this post we will see how to replace text in a Pandas. The short answer of this questions is: **(1) Replace character in Pandas column** ```python df['Depth'].str.replace('.',',') ``` **(2) Replace text in the whole Pandas DataFrame** ```python df.replace('\.',',', regex=True) ``` We will see several practical examples on how to replace text in Pandas columns and DataFrames. Suppose we have DataFrame like: | Date | Time | Latitude | Longitude | Depth | Magnitude Type | | ---------- | -------- | -------- | ---------- | ----- | -------------- | | 12/27/2016 | 23:20:56 | 45.7192 | 26.5230 | 97.0 | MWW | | 12/28/2016 | 08:18:01 | 38.3754 | \-118.8977 | 10.8 | ML | | 12/28/2016 | 08:22:12 | 38.3917 | \-118.8941 | 12.3 | ML | | 12/28/2016 | 09:13:47 | 38.3777 | \-118.8957 | 8.8 | ML | | 12/28/2016 | 12:38:51 | 36.9179 | 140.4262 | 10.0 | MWW | ## Replace single character in Pandas Column with .str.replace First let's start with the most simple example - replacing a single character in a single column. We are going to use the string method - `replace`: ```python df['Depth'].str.replace('.',',') ``` A warning message might be shown - for this one you can check the section below: > FutureWarning: The default value of regex will change from True to False in a future version. If you are using regex than you can specify it by: ```python df['Depth'].str.replace('.',',', regex= True) ``` result: ``` 23405 97 23406 10,8 23407 12,3 ``` ## Replace regex pattern in Pandas Column Let say that you would like to use a **regex in order to replace specific text patterns in Pandas.** For example let's change the date format of you Pandas DataFrame from: - mm/dd/yyyy to - yyyy-mm-dd This can be done by using regex flag and using regex groups like: ```python df['Date'].str.replace(r'(\d{2})/(\d{2})/(\d{4})', r"\3-\2-\1", regex=True) ``` result: ``` 23405 2016-27-12 23406 2016-28-12 23407 2016-28-12 ``` **How does the regex replace in Pandas work?** . So this one is translated as `(\d{2})`: - 1st capturing group `(\d{2})` - `\d` \- matches a digit (equivalent to \[0-9\]) - {2} matches the previous token exactly 2 times Then in the replacement part we are replaced by the group numbers - 1st group is `\1`. So you can try a simple exercise - to change the format to: dd mm 'yy . The answer is below: ```python df['Date'].str.replace(r'(\d{2})/(\d{2})/\d{2}(\d{2})', r"\2 \1 '\3", regex=True) ``` result: ``` 23405 27 12 '16 23406 28 12 '16 23407 28 12 '16 ``` ## Replace text in whole DataFrame If you like to replace values in all columns in your Pandas DataFrame then you can use syntax like: ```python df.replace('\.',',', regex=True) ``` If you don't specify the columns then the replace operation will be done over all columns and rows. ## Replace text with conditions in Pandas with lambda and .apply/.applymap **`.applymap` is another option to replace text and string in Pandas.** This one is useful if you have additional conditions - sometimes the regex might be too complex or to not work. In this case the **replacement can be done by lambda and `.apply` for columns**: ```python df['Date'].apply(lambda x: x.replace('/', '-')) ``` result: | Date | Time | Latitude | Longitude | Depth | Magnitude Type | | ---------- | -------- | -------- | ---------- | ----- | -------------- | | 12/27/2016 | 23:20:56 | 45.7192 | 26.523 | 97 | MWW | | 12/28/2016 | 08:18:01 | 38.3754 | \-118.8977 | 10.8 | ML | and **`.applymap` for the whole DataFrame**: ```python df.applymap(lambda x: x.replace('/', '-')) ``` result: | Date | Time | Latitude | Longitude | Depth | Magnitude Type | | ---------- | -------- | -------- | ---------- | ----- | -------------- | | 12/27/2016 | 23:20:56 | 45,7192 | 26,523 | 97 | MWW | | 12/28/2016 | 08:18:01 | 38,3754 | \-118,8977 | 10,8 | ML | Note that this solution might be slower than the others. ### Conditional replace in Pandas If you like to apply condition to your replacement in Pandas you can use syntax like: ```python df['Date'].apply(lambda x: x.replace('/', '-') if '20' in x else x) ``` result: ``` 23405 12-27-2016 23406 12-28-2016 23407 12-28-2016 ``` ## Warning: FutureWarning: The default value of regex will change from True to False in a future version. If you get an warning like: > FutureWarning: The default value of regex will change from True to False in a future version. It means that you will need to explicitly set the `regex` parameter for `replace` method: ```python df.replace('\.',',', regex=True) ``` The official documentation of method `replace` is: ``` `regex` -** bool, default True** Determines if the passed-in pattern is a regular expression: * If True, assumes the passed-in pattern is a regular expression. * If False, treats the pattern as a literal string Cannot be set to False if pat is a compiled regex or repl is a callable. ``` So the difference is how the passed pattern will be parsed - as a regex expression or not. So in a practical example: ```python df.replace('.',',') ``` this will not change the values of the DataFrame: ``` 45.7192 ``` While the regex expression - `'\.'` and `regex=True`: ```python df.replace('\.',',', regex=True) ``` will change them: ``` 45,7192 ``` ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/replace/replace-text-pandas-dataframe-column.ipynb?ref=datascientyst.com) - [pandas.Series.str.replace](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.replace.html?ref=datascientyst.com) - [pandas.Series.replace](https://pandas.pydata.org/docs/reference/api/pandas.Series.replace.html?ref=datascientyst.com) - [pandas.Series.apply](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.apply.html?ref=datascientyst.com) - [pandas.DataFrame.applymap](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.applymap.html?ref=datascientyst.com) ### How to Remove Timezone from a DateTime Column in Pandas URL: https://datascientyst.com/remove-timezone-datetime-column-pandas/ Last updated: 2022-02-25T21:08:24.000Z In this tutorial, we'll look at how to **remove the timezone info from a datetime column in a Pandas DataFrame.** The short answer of this questions is: (1) Use method `.dt.tz_localize(None)` ```python df['time_tz'].dt.tz_localize(None) ``` (2) With `.dt.tz_convert(None)` ```python df['time_tz'].dt.tz_convert(None) ``` This will help you when you need to compare with another column which doesn't have timezone info. Otherwise you will get an errors like: - **`TypeError: Timestamp subtraction must have the same timezones or no timezones`** - **`TypeError: Invalid comparison between dtype=datetime64[ns] and DatetimeArray`** Let's show simple example on **removing the timezone information in Pandas.** Starting with creating DataFrame like: ```python import pandas as pd import datetime import pytz dates = ['2021-08-01', '2021-08-02', '2021-08-03'] timestamps = [ datetime.datetime(2021, 8, 1, 12, 30, 41, 775854, tzinfo=pytz.timezone('US/Pacific')), datetime.datetime(2021, 8, 2, 12, 31, 12, 432523, tzinfo=pytz.timezone('US/Pacific')), datetime.datetime(2021, 8, 3, 12, 29, 59, 123512, tzinfo=pytz.timezone('US/Pacific')), ] df = pd.DataFrame({'start_date': dates, 'time': timestamps}) df ``` Data: | start\_date | time | time\_tz | | ----------- | -------------------------- | -------------------------------- | | 2021-08-01 | 2021-08-01 12:30:41.775854 | 2021-08-01 13:23:41.775854-07:00 | | 2021-08-02 | 2021-08-02 12:31:12.432523 | 2021-08-02 13:24:12.432523-07:00 | | 2021-08-03 | 2021-08-03 12:29:59.123512 | 2021-08-03 13:22:59.123512-07:00 | Now let's check the data stored in column `time`: ``` 0 2021-08-01 13:23:41.775854-07:00 1 2021-08-02 13:24:12.432523-07:00 2 2021-08-03 13:22:59.123512-07:00 Name: time, dtype: datetime64[ns, US/Pacific] ``` We have `dtype`: `datetime64[ns, US/Pacific]` If you like to compare this information with another datetime column without the timezone info you will get an errors like: - **`TypeError: Timestamp subtraction must have the same timezones or no timezones`** - **`TypeError: Invalid comparison between dtype=datetime64[ns] and DatetimeArray`** If you compare the column without the timezone info you will not face an error. ```python df['time'] > pd.to_datetime(df['start_date']) ``` working with timezone info is causing the errors mentioned above. ## Remove TimeZone from DateTime column in Pandas ### .dt.tz\_localize(None) In order to drop the timezone info from this column you can use: ```python df['time_tz'].dt.tz_localize(None) ``` which will result into: ``` 0 2021-08-01 13:23:41.775854 1 2021-08-02 13:24:12.432523 2 2021-08-03 13:22:59.123512 Name: time_tz, dtype: datetime64[ns] ``` Now you can use this column without getting the errors mentioned above. ### .dt.tz\_convert(None) Another option to deal with TimeZone info is by using the method: `.dt.tz_convert('UTC')`. You can delete or add the timezone info: ```python df['time_tz'].dt.tz_convert(None) ``` or ```python df['time_tz'].dt.tz_convert('UTC') ``` ![](https://datascientyst.com/content/images/2021/11/remove-timezone-datetime-column-pandas.png) ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/datetime/remove-timezone-datetime-column-pandas.ipynb?ref=datascientyst.com) - [pandas.Series.dt.tz\_convert](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.tz%5Fconvert.html?ref=datascientyst.com) - [pandas.Series.tz\_localize](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.tz%5Flocalize.html?ref=datascientyst.com) ### How to Convert DataFrame to List of Dictionaries in Pandas URL: https://datascientyst.com/convert-dataframe-list-dictionaries-pandas/ Last updated: 2021-10-29T11:55:25.000Z In this quick tutorial, we'll cover how to **convert Pandas DataFrame to a list of dictionaries.** Below you can find the quick **answer of DataFrame to list of dictionaries:** ```python df.to_dict('records') ``` Let's explain the solution in a practical example. Suppose we have DataFrame with data like: ```python import pandas as pd df = pd.read_csv('https://raw.githubusercontent.com/softhints/Pandas-Tutorials/master/data/population/population.csv') df ``` | Region | 1500 | 1600 | 1700 | 1750 | 2050 | 2150 | | -------------------------- | ---- | ---- | ---- | ---- | ---- | ---- | | World | 585 | 660 | 710 | 791 | 9725 | 9746 | | Africa | 86 | 114 | 106 | 106 | 2478 | 2308 | | Asia | 282 | 350 | 411 | 502 | 5267 | 5561 | | Europe | 168 | 170 | 178 | 190 | 734 | 517 | | Latin America \[Note 1\] ​ | 40 | 20 | 10 | 16 | 784 | 912 | ## Convert whole DataFrame to list of dictionaries Which you would like to convert to list of dictionaries like: ``` {'Region': 'World', '1500': 585, '1600': 660, '1700': 710, '1750': 791, '1800': 978, '1850': 1262, '1900': 1650, '1950': 2521, '1999': 6008, '2008': 6707, '2010': 6896, '2012': 7052, '2050': 9725, '2150': 9746}, {'Region': 'Africa', '1500': 86, '1600': 114, '1700': 106, ``` In order to achieve this behaviour you can use different approaches. First approach will use Pandas method `to_dict('records')`: ```python df.to_dict('records') ``` This method have several possible options: - 'dict’ (default) : dict like {column -> {index -> value}} - 'list’ : dict like {column -> \[values\]} - 'series’ : dict like {column -> Series(values)} - 'split’ : dict like {'index’ -> \[index\], 'columns’ -> \[columns\], 'data’ -> \[values\]} - 'tight’ : dict like {'index’ -> \[index\], 'columns’ -> \[columns\], 'data’ -> \[values\], 'index\_names’ -> \[index.names\], 'column\_names’ -> \[column.names\]} - 'records’ : list like \[{column -> value}, … , {column -> value}\] - 'index’ : dict like {index -> {column -> value}} ## Convert columns to list of dictionaries If you want to convert only some columns to a list of dictionaries you can use similar syntax. This can be achieved by selecting the columns and then applying method `.to_dict('records')`: ```python df[['Region', '1500', '1600', '1700']].to_dict('records') ``` Result is subset of the selected columns: ``` [{'Region': 'World', '1500': 585, '1600': 660, '1700': 710}, {'Region': 'Africa', '1500': 86, '1600': 114, '1700': 106}, {'Region': 'Asia', '1500': 282, '1600': 350, '1700': 411}, {'Region': 'Europe', '1500': 168, '1600': 170, '1700': 178}, {'Region': 'Latin America [Note 1] \u200b', '1500': 40, '1600': 20, '1700': 10}, {'Region': 'Northern America [Note 1] \u200b', '1500': 6, '1600': 3, '1700': 2}, {'Region': 'Oceania', '1500': 3, '1600': 3, '1700': 3}] ``` Transformation is visible from the image below: ![](https://datascientyst.com/content/images/2021/10/convert-dataframe-list-dictionaries-pandas.png) ## Convert DataFrame to list of dictionaries - column wise What if you like to get list of dictionaries column wise like: ``` {'Region': {0: 'World', 1: 'Africa', 2: 'Asia', 3: 'Europe', 4: 'Latin America [Note 1] \u200b', 5: 'Northern America [Note 1] \u200b', 6: 'Oceania'}, '1500': {0: 585, 1: 86, 2: 282, 3: 168, 4: 40, 5: 6, 6: 3}, '1600': {0: 660, 1: 114, 2: 350, 3: 170, 4: 20, 5: 3, 6: 3}, '1700': {0: 710, 1: 106, 2: 411, 3: 178, 4: 10, 5: 2, 6: 3}, ``` You can achieve this transformation by transposing the DataFrame with `.T` and using option `index`: ```python df.T.to_dict('index') ``` ## Custom Conversion of DataFrame to list of dictionaries Finally lets cover the case when there is a custom logic for **transformation of DataFrame to list of dictionaries in Pandas.** This option is a bit slower than the rest and works fine for small and medium sized data. We are going to iterate over all rows by: ```python data_dict = [] for index, row in df[['Region', '1500', '1600', '1700']].iterrows(): data_dict.append({ 'Region': row['Region'], '1500': row['1500'], '1600': row['1600'], '1700': row['1700'], }) ``` the result would be: ``` [{'Region': 'World', '1500': 585, '1600': 660, '1700': 710}, {'Region': 'Africa', '1500': 86, '1600': 114, '1700': 106}, {'Region': 'Asia', '1500': 282, '1600': 350, '1700': 411}, {'Region': 'Europe', '1500': 168, '1600': 170, '1700': 178}, {'Region': 'Latin America [Note 1] \u200b', '1500': 40, '1600': 20, '1700': 10}, {'Region': 'Northern America [Note 1] \u200b', '1500': 6, '1600': 3, '1700': 2}, {'Region': 'Oceania', '1500': 3, '1600': 3, '1700': 3}] ``` To convert all columns to list of dicts with custom logic you can use code like: ```python data_dict = [] for index, row in df[df.columns].iterrows(): data_dict.append({ row.to_list()[0] : row.to_list()[1:] }) data_dict ``` result: ``` [{'World': [585, 660, 710, 791, 978, 1262, 1650, 2521, 6008, 6707, 6896, 7052, 9725, 9746]}, {'Africa': [86, 114, 106, 106, 107, ``` ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/convert/convert-dataframe-list-dictionaries-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.to\_dict](https://pandas.pydata.org/pandas-docs/dev/reference/api/pandas.DataFrame.to%5Fdict.html?ref=datascientyst.com) ### Opposite of Melt in Python and Pandas URL: https://datascientyst.com/opposite-of-melt-python-pandas/ Last updated: 2021-12-02T10:48:10.000Z In this short guide, you'll see what is the **opposite operation of melt in Pandas and Python**. You can find a useful example. The short answer of the question above is: ```python df_m.pivot(*df_m).reset_index() ``` Let's show a detailed example on the above: ```python import pandas as pd df_pop = pd.read_csv('https://raw.githubusercontent.com/softhints/Pandas-Tutorials/master/data/population/population.csv') df_pop ``` | Region | 1500 | 1600 | 2010 | 2012 | 2050 | 2150 | | ------------------------ | ---- | ---- | ---- | ---- | ---- | ---- | | World | 585 | 660 | 6896 | 7052 | 9725 | 9746 | | Africa | 86 | 114 | 1022 | 1052 | 2478 | 2308 | | Asia | 282 | 350 | 4164 | 4250 | 5267 | 5561 | | Europe | 168 | 170 | 738 | 740 | 734 | 517 | | Latin America \[Note 1\] | 40 | 20 | 590 | 603 | 784 | 912 | In order to demo the opposite of melt operation let's perform melt on the above data: ```python df_m = df_pop.melt(id_vars=['Region']) df_m ``` The DataFrame `df_m` has the next data inside: | Region | variable | value | | --------------------------- | -------- | ----- | | World | 1500 | 585 | | Africa | 1500 | 86 | | Asia | 1500 | 282 | | Europe | 1500 | 168 | | Latin America \[Note 1\] | 1500 | 40 | | Northern America \[Note 1\] | 1500 | 6 | | Oceania | 1500 | 3 | | World | 1600 | 660 | | Africa | 1600 | 114 | | Asia | 1600 | 350 | So the years which were stored as columns after the melt operation are transformed to rows. Instead of the initial columns: ``` Index(['Region', '1500', '1600', '1700', '1750', '1800', '1850', '1900', '1950', '1999', '2008', '2010', '2012', '2050', '2150'], dtype='object') ``` after melt operation we end with 3 columns: - Region - the column on which we do the melt operation - variable - which is the column name of the old DataFrame - value - the corresponding value of the first DataFrame ## Reverse Melt Operation in Python and Pandas Now let's reverse the melt which was performed above. We are going to work with DataFrame df\_m. There are **several ways of reversing melt operation in Pandas.** In this post we will demonstrate the one which uses method `pivot` and `reset_index`: ```python df_m.pivot(*df_m).reset_index() ``` If you like to get the original DataFrame from you will need to rename the columns by `.rename_axis(None, axis='columns')`. So the full code will become: ```python df_m.pivot(*df_m).reset_index().rename_axis(None, axis='columns') ``` The difference between those two is the name of the index. Without `.rename_axis(None, axis='columns')` we will get: > Index(\['Region', '1500', '1600', '1700', '1750', '1800', '1850', '1900', > '1950', '1999', '2008', '2010', '2012', '2050', '2150'\], > dtype='object', name='variable') with `.rename_axis(None, axis='columns')` we will have different index name after the pivot.: > Index(\['Region', '1500', '1600', '1700', '1750', '1800', '1850', '1900', > '1950', '1999', '2008', '2010', '2012', '2050', '2150'\], > dtype='object') ![](https://datascientyst.com/content/images/2021/10/opposite-of-melt-python-pandas.png) Note: Please note that the index is sorted alphabetically and doesn't match the original sort. ## Opposite of melt on few values only Finally let's see the example when you like to melt or reverse it on a few variables. This can be done by using parameter `value_vars` of method `melt`: ```python df_m = df_pop.melt(id_vars=['Region'], value_vars=['1500', '1600', '1700']) df_m ``` result: | Region | variable | value | | -------------------------- | -------- | ----- | | World | 1500 | 585 | | Africa | 1500 | 86 | | Asia | 1500 | 282 | | Europe | 1500 | 168 | | Latin America \[Note 1\] ​ | 1500 | 40 | The reverse operation have additional parameter `column`: ```python df_m.pivot(index='Region', columns='variable')['value'] ``` | variable | 1500 | 1600 | 1700 | | --------------------------- | ---- | ---- | ---- | | Region | | | | | Africa | 86 | 114 | 106 | | Asia | 282 | 350 | 411 | | Europe | 168 | 170 | 178 | | Latin America \[Note 1\] | 40 | 20 | 10 | | Northern America \[Note 1\] | 6 | 3 | 2 | ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/split/split-pandas-list-column-into-multiple-columns.ipynb?ref=datascientyst.com) - [pandas.melt](https://pandas.pydata.org/docs/reference/api/pandas.melt.html?ref=datascientyst.com) - [pandas.pivot](https://pandas.pydata.org/docs/reference/api/pandas.pivot.html?ref=datascientyst.com) ### How to Split Column into Multiple Columns in Pandas URL: https://datascientyst.com/split-pandas-list-column-into-multiple-columns/ Last updated: 2022-02-19T07:44:44.000Z ## 1\. Overview Here are two approaches to **split a column into multiple columns in Pandas:** - list column - string column separated by a delimiter. Below we can find both examples: (1) **Split column (list values) into multiple columns** ```python pd.DataFrame(df["langs"].to_list(), columns=['prim_lang', 'sec_lang']) ``` (2) **Split column by delimiter in Pandas** ```python pd.DataFrame(df["skills"].str.split(',').fillna('[]').tolist()) ``` In **Pandas to split column we can use method `.str.split(',')`** \- which takes a delimiter as a parameter. Next we will see how to apply both ways into practical examples. ## 2\. Split List Column into Multiple Columns For the first example we will create a simple DataFrame with 1 column which stores a list of two languages. We are going to generate 10 random lists of subset of languages: ```python import random langs = ['Python', 'Java' , 'JS', 'C', 'C+'] df = pd.DataFrame({"langs": [ [random.choice(langs) for i in range(0,2)] for _ in range(10)]}) ``` Our DataFrame looks like this: | langs | | -------------- | | \[C+, C+\] | | \[Python, JS\] | | \[C+, C+\] | | \[Java, Java\] | | \[Java, C\] | In order to **split this single column(which contain list values) into two columns** we will use the next syntax: ```python pd.DataFrame(df["langs"].to_list(), columns=['prim_lang', 'sec_lang']) ``` The result of the split is: | prim\_lang | sec\_lang | | ---------- | --------- | | C+ | C+ | | Python | JS | | C+ | C+ | | Java | Java | | Java | C | How does it work? The method `df["langs"].to_list()` is converting the initial column into list of lists: ``` [['C+', 'C+'], ['Python', 'JS'], ['C+', 'C+'], ['Java', 'Java'], ['Java', 'C'], ['Java', 'C+'], ['Python', 'C'], ['Python', 'C+'], ['C', 'Java'], ['C+', 'JS']] ``` Note: This method will work only if the stored values are lists. If you have string values separated by columns check Example 2. ![](https://datascientyst.com/content/images/2021/10/split-pandas-list-column-into-multiple-columns.png) ## 3\. Split Column by delimiter in Pandas Now let's say that instead of storing lists like: `['C+', 'C+']` you have only the values separated by delimiter. In this case is a comma like `'C+', 'C+'`. Lets have data like the one below: | skills | internship | location | | -------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | -------------- | | Software Testing | Software Testing | Noida | | Java, SQL, Unix, Oracle, MS SQL Server, Hibernate (Java), Shell Scripting, Spring MVC, REST API | Technical Operations - Networking And Monitoring | Noida | | English Proficiency (Spoken), English Proficiency (Written), Hindi Proficiency (Spoken), Hindi Proficiency (Written) | Software Project Management | Work From Home | | HTML, CSS, Flask, Python, Django | Web Development | Work From Home | | HTML, CSS, JavaScript, ReactJS, Redux | Front End Development | Work From Home | And we would like to **split the column skills by delimiter into multiple columns**. This time the number of elements is not fixed! We can use Pandas string method `.str.split(',')` in order to split the values into lists of lists. If you have missing data you need to ensure that you default it by empty list by `.fillna('[]')`: ```python pd.DataFrame(df["skills"].str.split(',').fillna('[]').tolist()) ``` This will create DataFrame like: | 0 | 1 | 2 | 44 | 45 | | ---------------------------- | ----------------------------- | -------------------------- | ---- | ---- | | Software Testing | None | None | None | None | | Java | SQL | Unix | None | None | | English Proficiency (Spoken) | English Proficiency (Written) | Hindi Proficiency (Spoken) | None | None | | HTML | CSS | Flask | None | None | | HTML | CSS | JavaScript | None | None | As you can see the result DataFrame has 45 columns. Which means that one of the rows has 45 values separated by comma. In order to find which row(s) have most values we can use syntax like - test the last column for all non null elements: ```python pd.DataFrame(df["skills"].str.split(',').fillna('[]').tolist())[45].dropna() ``` And we get as output: ``` 1242 Gitlab ``` ## 4\. Conclusion and Resources In this guide we saw how to split columns depending on the values which contain. We covered a column which contains lists and also splitting values separated by delimiter. Below you can find some useful resources: - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/split/split-pandas-list-column-into-multiple-columns.ipynb?ref=datascientyst.com) - [pandas.Series.str.split](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.split.html?ref=datascientyst.com) - [pandas.Series.tolist](https://pandas.pydata.org/docs/reference/api/pandas.Series.tolist.html?ref=datascientyst.com) ### How to apply function to single column in Pandas URL: https://datascientyst.com/apply-function-to-single-column-in-pandas/ Last updated: 2021-09-11T21:24:00.000Z In this quick tutorial, we'll cover **how to apply function to a single column in Pandas**. Here are two ways to apply function to column in DataFrame: **(1) Apply user defined function on column** ```python df['col'].map(my_function) ``` **(2) Apply lambda to function** ```python df['col'].apply(lambda x: x + 1) ``` For multiple columns check: [How to apply function to multiple columns in Pandas](https://datascientyst.com/apply-function-multiple-columns-pandas/). In the next section we will cover several different use cases and important details on the topic. Let's say that we have the following DataFrame: ```python import pandas as pd import numpy as np df = pd.DataFrame(np.random.randint(0,5,size=(5, 4)), columns=list('ABCD')) df ``` DataFrame: | A | B | C | D | | - | - | - | - | | 4 | 2 | 2 | 3 | | 1 | 0 | 0 | 3 | | 0 | 2 | 1 | 0 | | 3 | 2 | 4 | 2 | | 1 | 1 | 4 | 1 | × **Pro Tip!** For large DataFrames use Dask or swifter - check Option 4! ## Option 1: Pandas apply function to column The first example will show how to **define a function and then apply it on a column from a Pandas DataFrame**. First we will define a function which will be applied on the column by method - `pd.apply`. Then we will called that function for column `A`: ```python def my_function(x): return x ** 2 df['A'].apply(my_function) ``` The result is squared values for each cell: ``` 0 16 1 1 2 0 3 9 4 1 Name: A, dtype: int64 ``` ## Option 2: Pandas apply function to column by `map` \*\*A better way to apply function to a single column is by using Pandas `map` method. \*\* Why is it better? Because `apply` is designed for multiple columns while `map` is intended for Pandas Series. A single column from Pandas is equal to a Pandas Series or 1 dimensional array. Method `map` can be slightly faster than `apply` for large DataFrames. So the apply function by map can be done by: ```python def my_function(x): return x ** 2 df['A'].map(my_function) ``` The result is the same as Option 1. ## Option 3: Pandas apply anonymous function / lambda to column Sometimes a **lambda or anonymous function** is what you would like to **apply to a column**. The syntax is very simple: ```python df['A'].map(lambda x: x ** 2) ``` or: ```python df['A'].apply(lambda x: x ** 2) ``` The difference is the same: `apply` method will be applied on DataFrame level while `map` is applied on Series level. So you can do: ```python df[['A', 'B']].apply(lambda x: x ** 2) ``` in order to apply lambda on multiple functions. × **Pro Tip!** It's recommended to show intention of what you would like to do: 1) use map for single columns or 2) use apply if in future other columns should be added! ## Option 4: Speed up Pandas apply function to column Finally let's cover how to **speed up applying function to single column in Pandas.** To optimize execution there are several options like: - Dask - Swifter You can find more info in the Resource section. To test all options we will create DataFrame with shape - `10000 rows × 4 columns`. **Dask - apply function to column** \- the fastest tested way: ```python import dask.dataframe as dd ddf = dd.from_pandas(df, npartitions=2) ddf["A"].apply(fnc, meta=('A', 'int64')) ``` The result is: ``` 612 µs ± 3.31 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) ``` **swifter - fast apply function to column** \- it's much faster the method `apply` and a bit slower than Dask - in the tested example: ```python df['A'].swifter.apply(fnc) ``` result: ``` 752 µs ± 3.39 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) ``` **Pandas apply** ``` 3.63 ms ± 28.2 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) ``` **Pandas map** ``` 3.57 ms ± 25 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) ``` ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/apply/apply-function-to-single-column-in-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.apply](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.apply.html?ref=datascientyst.com) - [pandas.Series.map](https://pandas.pydata.org/docs/reference/api/pandas.Series.map.html?highlight=map&ref=datascientyst.com#pandas.Series.map) - [pandas.Series.swifter.apply](https://github.com/jmcarpenter2/swifter/blob/master/docs/documentation.md?ref=datascientyst.com) - [dask.dataframe.DataFrame.apply](https://docs.dask.org/en/latest/generated/dask.dataframe.DataFrame.apply.html?ref=datascientyst.com) - [How to Create a Pandas DataFrame of Random Integers](https://datascientyst.com/how-to-create-a-dataframe-of-random-integers-with-pandas/) ### How to Merge Two DataFrames on Index in Pandas URL: https://datascientyst.com/merge-two-dataframes-on-index-pandas/ Last updated: 2021-09-10T22:30:52.000Z In this short tutorial, we'll show how to **merge two DataFrames on index in Pandas.** Here are two approaches to merge DataFrames on index: **(1) Use method `merge` with `left_index` and `right_index`** ```python pd.merge(df1, df2, left_index=True, right_index=True) ``` **(2) Method `concat` with `axis=1`** ```python pd.concat([df1, df2], axis=1) ``` To start, let's say that we have two DataFrames: **df1** | A | B | C | D | | -- | -- | -- | -- | | A0 | B0 | C0 | D0 | | A1 | B1 | C1 | D1 | | A2 | B2 | C2 | D2 | | A3 | B3 | C3 | D3 | **df2** | A | B | C | D | | -- | -- | -- | -- | | A4 | B4 | C4 | D4 | | A5 | B5 | C5 | D5 | | A6 | B6 | C6 | D6 | | A7 | B7 | C7 | D7 | ## Option 1: Pandas: merge on index by method `merge` The first example will show **how to use method `merge` in combination with `left_index` and `right_index`**. This will merge on index by `inner join` \- only the rows from the both DataFrames with similar index will be added to the result: ```python pd.merge(df1, df2, left_index=True, right_index=True) ``` The result is: | A\_x | B\_x | C\_x | D\_x | A\_y | B\_y | C\_y | D\_y | | ---- | ---- | ---- | ---- | ---- | ---- | ---- | ---- | | A0 | B0 | C0 | D0 | A4 | B4 | C4 | D4 | | A1 | B1 | C1 | D1 | A5 | B5 | C5 | D5 | | A2 | B2 | C2 | D2 | A6 | B6 | C6 | D6 | | A3 | B3 | C3 | D3 | A7 | B7 | C7 | D7 | ![merge-two-dataframes-on-index-pandas](https://datascientyst.com/content/images/2021/09/merge-two-dataframes-on-index-pandas-1.png) As you can see the result DataFrame is a combination of both merged on the index. The column names have suffixes: `_x` and `_y`. In order to change the suffixes use the following syntax: ```python df_m = pd.merge(df1, df2, left_index=True, right_index=True, suffixes=('_left', '_right')) df_m ``` which uses parameter `suffixes: 'Suffixes' = ('_x', '_y')` The functions is described as: > Merge DataFrame or named Series objects with a database-style join. ### Duplicated index while merging on index In case of duplicated index while merging on index the rows will be duplicated: ![merge-two-dataframes-on-index-pandas-duplicated-indexes](https://datascientyst.com/content/images/2021/09/merge-two-dataframes-on-index-pandas-duplicated-indexes.png) ## Option 2: Pandas: merge on index by `concat` and axis=1 Alternative **solution is to merge both DataFrames by using method `concat`**. For this approach you need to set the parameter `axis=1` which is equivalent to `axis='columns'`. ```python pd.concat([df1, df2], axis='columns') ``` The result is similar to the previous one with differences in the column names. Using `concat` will not change the column names. | A | B | C | D | A | B | C | D | | -- | -- | -- | -- | -- | -- | -- | -- | | A0 | B0 | C0 | D0 | A4 | B4 | C4 | D4 | | A1 | B1 | C1 | D1 | A5 | B5 | C5 | D5 | | A2 | B2 | C2 | D2 | A6 | B6 | C6 | D6 | | A3 | B3 | C3 | D3 | A7 | B7 | C7 | D7 | As you can see we have duplicate names for the columns. In order to avoid `concat` with duplicated columns use `verify_integrity`: ```python df_m = pd.concat([df1, df2], axis='columns', verify_integrity=True) df_m ``` This will raise error: > ValueError: Indexes have overlapping values: Index(\['A', 'B', 'C', 'D'\], dtype='object') If you still prefer to have duplicate columns you can access them in the normal way: ```python df_m['A'] ``` result: | A | A | | -- | -- | | A0 | A4 | | A1 | A5 | | A2 | A6 | | A3 | A7 | ## Option 3: Pandas: merge on index - `join` Method **`join` can be used for the merge operation**. By default is `left join`. In case of duplicated column names will error: > ValueError: columns overlap but no suffix specified: Index(\['A', 'B', 'C', 'D'\], dtype='object') In order to avoid errors use the following syntax: ```python df1.join(df2, lsuffix='_x') ``` ## Conclusion: Pandas: merge on index - `merge` vs `concat` Finally let's compare the both ways: - `merge` will assign suffixes for both DataFrames / `concat` will not change the column names - `merge` has control on the merge operation. By default the operation is `inner join` \- you can change it to `left`, `right`, `outer` etc - `merge` So if we add extra row for **df2** as: | | A | B | C | D | | - | -- | -- | -- | -- | | 0 | A4 | B4 | C4 | D4 | | 1 | A5 | B5 | C5 | D5 | | 2 | A6 | B6 | C6 | D6 | | 3 | A7 | B7 | C7 | D7 | | 4 | A8 | B8 | C8 | D8 | Then concat will do full outer join by default: | | A | B | C | D | A | B | C | D | | - | --- | --- | --- | --- | -- | -- | -- | -- | | 0 | A0 | B0 | C0 | D0 | A4 | B4 | C4 | D4 | | 1 | A1 | B1 | C1 | D1 | A5 | B5 | C5 | D5 | | 2 | A2 | B2 | C2 | D2 | A6 | B6 | C6 | D6 | | 3 | A3 | B3 | C3 | D3 | A7 | B7 | C7 | D7 | | 4 | NaN | NaN | NaN | NaN | A8 | B8 | C8 | D8 | While merge will do only inner join and the result will be the same as in Step 1\. In order to change `merge` behaviour you need to change the `how` parameter: ```python df_m = pd.merge(df1, df2, left_index=True, right_index=True, how='outer') df_m ``` So the **The best way to merge on an index is to use the method `merge`.** ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/merge/merge-two-dataframes-on-index-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.merge](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.merge.html?ref=datascientyst.com) - [pandas.concat](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.concat.html?highlight=concat&ref=datascientyst.com) - [Merge, join, concatenate and compare](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/merging.html?ref=datascientyst.com) ### Use read_csv to skip rows with condition based on values in Pandas URL: https://datascientyst.com/read_csv-skip-rows-condition-values-pandas/ Last updated: 2022-10-29T21:15:21.000Z In this tutorial, we'll look at how to **read CSV files by `read_csv` and skip rows with a conditional statement in Pandas**. In addition, we'll also see how to optimise the reading performance of the `read_csv` method with Dask. This option is useful if you face memory issues using `read_csv`. To start let's say that we have the following CSV file: | gender | race/ethnicity | math score | reading score | writing score | | ------ | -------------- | ---------- | ------------- | ------------- | | female | group B | 72 | 72 | 74 | | female | group C | 69 | 90 | 88 | | female | group B | 90 | 95 | 93 | | male | group A | 47 | 57 | 44 | | male | group C | 76 | 78 | 75 | ``` 1000 rows × 8 columns ``` ## Step 1: Read CSV file skip rows with query condition in Pandas By default Pandas `skiprows` parameter of method `read_csv` is supposed to filter rows based on row number and not the row content. So the default behavior is: ```python pd.read_csv(csv_file, skiprows=5) ``` The code above will result into: ``` 995 rows × 8 columns ``` But let's say that we would like to **skip rows based on the condition on their content**. This can be achieved by reading the CSV file in chunks with `chunksize`. The results will be filtered by `query` condition: ```python gen = pd.read_csv(csv_file, chunksize=10000000) df = pd.concat((x.query("lunch == 'standard'") for x in gen), ignore_index=True) ``` result: ``` 645 rows × 8 columns ``` The above code will filter CSV rows based on column `lunch`. It will return only rows containing `standard` to the output. ## Step 2: Read CSV file with condition value higher than threshold In this step we are going to **compare the row value in the rows against integer value**. If the value is equal or higher we will load the row in the CSV file. Difference with the previous step is: - the usage of `dtype` parameter which defines a schema for the columns read from the CSV file - how to use query with column which contains space - `math score` \- surrounded by \` ```python schema={ "math score": int } gen = pd.read_csv(csv_file, dtype=schema, chunksize=10000000) df = pd.concat((x.query("`math score` >= 75") for x in gen), ignore_index=True) ``` The code above will filter all rows which contain `math score` higher or equal to 75: ``` 295 rows × 8 columns ``` ## Step 3: Read CSV file and post filter condition in Pandas For small and medium CSV files it's fine to **read the whole file and do a post filtering based on read values**. This can be achieved by: ```python df = pd.read_csv(csv_file) df = df[df['lunch'] != 'standard'] ``` result: ``` 355 rows × 8 columns ``` So first we read the whole file. Next we are filtering the results based on one or multiple conditions. ## Step 4: Read CSV with conditional filtering by Dask Finally let's see how to read a CSV file with condition and optimised performance. To install Dask use: ```python pip install dask ``` Dask offers a lazy reader which can optimize performance of `read_csv`. So first we can read the CSV file, then apply the filtering and finally to compute the results: ```python import dask.dataframe as dd df = dd.read_csv(csv_file) df = df[(df['reading score'] >= 55) & (df['reading score'] <= 75)] df = df.compute() ``` result is: ``` 496 rows × 8 columns ``` ![read_csv-skip-rows-condition-values-pandas](https://datascientyst.com/content/images/2022/10/read_csv-skip-rows-condition-values-pandas.webp) ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/read%5Fcsv/pandas-read-csv-file-read%5Fcsv-skiprows.ipynb?ref=datascientyst.com) - [pandas.DataFrame.query](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.query.html?ref=datascientyst.com) - [pandas.read\_csv](https://pandas.pydata.org/pandas-docs/dev/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) - [Feature Request: "Skiprows" by a condition or set of conditions](https://github.com/pandas-dev/pandas/issues/32072?ref=datascientyst.com) - [dask](https://pypi.org/project/dask/?ref=datascientyst.com) - [How to Skip First Rows in Pandas read\_csv and skiprows?](https://datascientyst.com/pandas-read-csv-file-read%5Fcsv-skiprows/) ### How to Skip First Rows in Pandas read_csv and skiprows? URL: https://datascientyst.com/pandas-read-csv-file-read_csv-skiprows/ Last updated: 2021-09-08T21:13:42.000Z Do you need to s**kip rows while reading CSV file with read\_csv in Pandas**? If so, this article will show you how to skip first rows of reading file. Method `read_csv` has parameter `skiprows` which can be used as follows: **(1) Skip first rows reading CSV file in Pandas** ```python pd.read_csv(csv_file, skiprows=3, header=None) ``` **(2) Skip rows by index with `read_csv`** ```python pd.read_csv(csv_file, skiprows=[0,2]) ``` Lets check several practical examples which will cover all aspects of **reading CSV file and skipping rows**. To start lets say that we have the next CSV file: ```bash !cat '../data/csv/multine_header.csv' ``` CSV file with multiple headers (to learn more about [reading a CSV file with multiple headers](https://datascientyst.com/read-excel-csv-multiple-line-headers-using-pandas/)): ``` Date,Company A,Company A,Company B,Company B ,Rank,Points,Rank,Points 2021-09-06,1,7.9,2,6 2021-09-07,1,8.5,2,7 2021-09-08,2,8,1,8.1 ``` ## Step 1: Skip first N rows while reading CSV file First example shows how to skip consecutive rows with Pandas `read_csv` method. There are 2 options: - skip rows in Pandas without using header - skip first N rows and use header for the DataFrame - check Step 2 In this Step Pandas read\_csv method will read data from row 4 (index of this row is 3). The newly created DataFrame will have autogenerated column names: ```python df = pd.read_csv(csv_file, skiprows=3, header=None) ``` This will result into: | 0 | 1 | 2 | 3 | 4 | | ---------- | - | --- | - | --- | | 2021-09-07 | 1 | 8.5 | 2 | 7.0 | | 2021-09-08 | 2 | 8.0 | 1 | 8.1 | ## Step 2: Skip first N rows and use header If parameter `header` of method `read_csv` is not provided than first row will be used as a header. In combination of parameters `header` and `skiprows` \- first the rows will be skipped and then first on of the remaining will be used as a header. In the example below 3 rows from the CSV file will be skipped. The forth one will be used as a header of the new DataFrame. ```python df = pd.read_csv(csv_file, skiprows=3) ``` | 2021-09-07 | 1 | 8.5 | 2 | 7 | | ---------- | - | --- | - | --- | | 2021-09-08 | 2 | 8 | 1 | 8.1 | ## Step 3: Pandas keep the header and skip first rows What if you need to keep the header and then the skip N rows? This can be achieved in several different ways. The most simple one is by builing a list of rows which to be skipped: ```python rows_to_skip = range(1,3) df = pd.read_csv(csv_file, skiprows=rows_to_skip) ``` result: | Date | Company A | Company A.1 | Company B | Company B.1 | | ---------- | --------- | ----------- | --------- | ----------- | | 2021-09-07 | 1 | 8.5 | 2 | 7.0 | | 2021-09-08 | 2 | 8.0 | 1 | 8.1 | As you can see **`read_csv` method keep the header and skip first 2 rows** after the header. ## Step 4: Skip non consecutive rows with `read_csv` by index Parameter `skiprows` is defined as: > Line numbers to skip (0-indexed) or number of lines to skip (int) at the start of the file. So to **skip rows 0 and 2 we can pass list of values to `skiprows`**: ```python df = pd.read_csv(csv_file, skiprows=[0,2]) ``` | Unnamed: 0 | Rank | Points | Rank.1 | Points.1 | | ---------- | ---- | ------ | ------ | -------- | | 2021-09-07 | 1 | 8.5 | 2 | 7.0 | | 2021-09-08 | 2 | 8.0 | 1 | 8.1 | ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/read%5Fcsv/pandas-read-csv-file-read%5Fcsv-skiprows.ipynb?ref=datascientyst.com) - [General parsing configuration - skiprows](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/io.html?ref=datascientyst.com#general-parsing-configuration) - [pandas.read\_csv](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read%5Fcsv.html?highlight=read%5Fcsv&ref=datascientyst.com) ### How to Rename Index in Pandas DataFrame URL: https://datascientyst.com/rename-index-in-pandas-dataframe/ Last updated: 2022-10-31T13:07:38.000Z There are two approaches to **rename index/columns in Pandas DataFrame**: **(1) Set new index name by `df.index.names`** ```python df.index.names = ['org_id'] ``` **(2) Rename index name with `rename_axis`** ```python df.rename_axis('org_id') ``` **(3) Rename column name with `rename_axis`** ```python df.rename_axis('col_index', axis=1) ``` ![](https://datascientyst.com/content/images/2022/10/rename-index-in-pandas-dataframe.webp) In the rest of this article you can find a few practical examples on index renaming for columns and rows. To start, let's create a sample DataFrame with 2 columns: ```python import pandas as pd df = pd.DataFrame({ 'name':['Softhints', 'DataScientyst'], 'url':['https://www.softhints.com', 'https://datascientyst.com'] }) ``` | | name | url | | - | ------------- | ------------------------- | | 0 | Softhints | https://www.softhints.com | | 1 | DataScientyst | https://datascientyst.com | ## Step 1: Check index/axis name in Pandas First let's check if there's an index name by: ```python df.index ``` The index don't have any name set: ```python RangeIndex(start=0, stop=2, step=1) ``` The same is for the columns: ```python df.columns ``` no name on column axis: ``` Index(['name', 'url'], dtype='object') ``` ## Step 2: Pandas rename index name - .rename\_axis() Pandas offers a method called `.rename_axis('companies')` which is intended for index renaming. To **rename index with `rename_axis` in Pandas DataFrame use**: ```python df.rename_axis('org_id') ``` this will result into: | | name | url | | ------- | ------------- | ------------------------- | | org\_id | | | | 0 | Softhints | https://www.softhints.com | | 1 | DataScientyst | https://datascientyst.com | You can notice the index name `org_id` which is on the top of the index. Let's check the index name by `df.index`: ``` RangeIndex(start=0, stop=2, step=1, name='org_id') ``` It's visible by `name='org_id'` For changing column index name you can use - `axis=1`: ```python df = df.rename_axis('company_data', axis=1) ``` result: | company\_data | name | url | | ------------- | ------------- | ------------------------- | | org\_id | | | | 0 | Softhints | https://www.softhints.com | | 1 | DataScientyst | https://datascientyst.com | and for columns we have `name` added to the index: ``` Index(['name', 'url'], dtype='object', name='company_data') ``` ## Step 3: Rename Pandas index with df.index.names Another option to rename any axis is to use the `names` attribute: `df.index.names`. So to rename the index name is by: ```python df.index.names = ['org_id'] ``` For columns we can use: ```python df.columns.names = ['company_data'] ``` The result is exactly the same as in the previous step. ## Step 4: Pandas rename index name -df.index.rename() Pandas offers method `index.rename` which can be used to change the index name for both rows and/or columns: ```python df.index = df.index.rename('test') ``` result: ``` RangeIndex(start=0, stop=2, step=1, name='test') ``` ## Step 5: Get Pandax index level by name Finally let's cover how index can be extracted by it's name: ```python df.index.get_level_values('test') ``` This is useful for hierarchical indexes. So to access data by index name use: ```python df.loc[df.index.get_level_values('test')] ``` ## Step 6: Pandas rename index column To rename column index in Pandas we need to pass parameter `axis=1` to method `rename_axis()`: ```python df.rename_axis('col_index', axis=1) ``` ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/index/rename-index-in-pandas-dataframe.ipynb?ref=datascientyst.com) - [pandas.DataFrame.rename\_axis](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.rename%5Faxis.html?ref=datascientyst.com) - [pandas.Index.name](https://pandas.pydata.org/docs/reference/api/pandas.Index.name.html?ref=datascientyst.com) - [pandas.Index.names](https://pandas.pydata.org/docs/reference/api/pandas.Index.names.html?ref=datascientyst.com) ### How to Create a Clickable Link(s) in Pandas DataFrame and JupyterLab URL: https://datascientyst.com/create-clickable-link-pandas-dataframe-jupyterlab/ Last updated: 2022-06-21T06:38:16.000Z Here are few different approaches to **create a hyperlink in Pandas DataFrame and JupyterLab**: (1) **Pandas method: `to_html` and parameter `render_links=True`**: ```python HTML(df.to_html(render_links=True, escape=False)) ``` (2) **Create clickable link from single column in Pandas DataFrame**: ```python def make_clickable(val): return f'{val}' df.style.format({'url': make_clickable}) ``` (3) **Create hyperlink from name and URL in Pandas with `apply`**: ```python df['name'] = df['name'].apply(lambda x: f'{x}') HTML(df.to_html(escape=False)) ``` In the next section, We'll review the steps to apply the above examples in practice. ## Step 1: Create a Hyperlink with `to_html` and `render_links` \- multiple columns The first example to cover will **create links from all columns which contain valid URL-s with option - `render_links`**. Parameter `render_links` is from boolean type and has default False. > Convert URLs to HTML links. So let's create a DataFrame with two columns: - name - url - ur2 ```python from IPython.display import HTML import pandas as pd df = pd.DataFrame({ 'name':['Softhints', 'DataScientyst'], 'url':['https://www.softhints.com', 'https://datascientyst.com'], 'url2':['https://www.blog.softhints.com/tag/pandas', 'https://datascientyst.com/tag/pandas'] }) ``` to render all links from this DataFrame we can use the next syntax: ```python HTML(df.to_html(render_links=True, escape=False)) ``` DataFrame contains hyperlinks: | | name | url | url2 | | - | ------------- | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | | 0 | Softhints | [https://www.softhints.com](https://www.softhints.com/?ref=datascientyst.com) | [https://www.blog.softhints.com/tag/pandas](https://www.blog.softhints.com/tag/pandas?ref=datascientyst.com) | | 1 | DataScientyst | [https://datascientyst.com](https://datascientyst.com/) | | As you can see from the table above all URLs are rendered as links. ## Step 2: Create a new hyperlink column as combination of others columns in Pandas In this example we can see how to create a method which is going to convert: - name - url from **two columns of Pandas DataFrame to a new column with a short hyperlink**. Let's say that our DataFrame has two values - name and url . ```python df = pd.DataFrame({ 'name':['Softhints', 'DataScientyst'], 'url':['https://www.softhints.com', 'https://datascientyst.com'] }) ``` To create a **new column which is a shortened hyperlink** we can make a function. Then we can use method `format` of `style` in order to apply the function to different columns - given as a dict: ```python def make_clickable(val): return f'{val}' df.style.format({'url': make_clickable}) ``` result: | | name | url | | - | ------------- | ----------------------------------------------------------------------------- | | 0 | Softhints | [https://www.softhints.com](https://www.softhints.com/?ref=datascientyst.com) | | 1 | DataScientyst | [https://datascientyst.com](https://datascientyst.com/) | To use only the URL without the name we can build another function: ```python def make_clickable(val): return f'{val}' df.style.format(make_clickable) ``` **Creating a new clickable column which is a combination from the other columns of the same DataFrame**: ```python df['link'] = df.apply(lambda x: make_clickable(x['url'], x['name']), axis=1) df.style ``` ## Step 3: Create hyperlink from URL using lambda and to\_html(escape=False) Finally, let's see how to **combine lambda and `to_html(escape=False)` to create a clickable link in DataFrame**. ```python from IPython.display import HTML df = pd.DataFrame({'name':['Pandas', 'Linux']}) df['name'] = df['name'].apply(lambda x: f'{x}') HTML(df.to_html(escape=False)) ``` As a result the column URL has clickable values: | | name | | - | -------------------------------------------------------------------- | | 0 | [Pandas](http://softhints.com/tutorial/Pandas?ref=datascientyst.com) | | 1 | [Linux](http://softhints.com/tutorial/Linux?ref=datascientyst.com) | ## Resources - [Notebook](https://github.com/softhints/Pandas-Tutorials/blob/master/styling/create-clickable-link-pandas-dataframe-jupyterlab.ipynb?ref=datascientyst.com) - [pandas.DataFrame.to\_html](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to%5Fhtml.html?ref=datascientyst.com) ### How to Fix - UnicodeDecodeError: invalid start byte - during read_csv in Pandas URL: https://datascientyst.com/pandas-read_csv-unicodedecodeerror-invalid-start-byte/ Last updated: 2024-01-05T13:08:16.000Z In this short guide, I'll show you\*\* how to solve the error: UnicodeDecodeError: invalid start byte while reading a CSV with Pandas\*\*: > pandas UnicodeDecodeError: 'utf-8' codec can't decode byte 0x97 in position 6785: invalid start byte The error might have several different reasons: - different encoding - bad symbols - corrupted file Below you can find quick solution for this error: **pandas UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte**. ```python df = pd.read_csv('file', encoding='utf-16') ``` Adding `encoding='utf-16'` to Pandas read\_csv() will solve the error. In the next steps you will find information on how to investigate and solve the error. As always all examples can be found in a handy: [Jupyter Notebook](https://github.com/softhints/datascientyst/blob/master/read%5Fcsv/pandas-read%5Fcsv-unicodedecodeerror-invalid-start-byte.ipynb?ref=datascientyst.com) ## 1: UnicodeDecodeError: invalid start byte while reading CSV file To start, let's demonstrate the error: UnicodeDecodeError while reading a sample CSV file with Pandas. The file content is shown below by Linux command `cat`: ``` ��a,b,c 1,2,3 ``` We can see some strange symbol at the file start: `��` Using method `read_csv` on the file above will raise error: ```python df = pd.read_csv('../data/csv/file_utf-16.csv') ``` raised error: > UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte ## 2: Solution of UnicodeDecodeError: change read\_csv encoding The first solution which can be applied in order to solve the error `UnicodeDecodeError` is to change the encoding for method `read_csv`. To use different encoding we can use parameter: `encoding`: ```python df = pd.read_csv('../data/csv/file_utf-16.csv', encoding='utf-16') ``` and the file will be read correctly. The weird start of the file was suggesting that probably the encoding is not utf-8. In order to check what is the correct encoding of the CSV file we can use next Linux command or Jupyter magic: ```python !file '../data/csv/file_utf-16.csv' ``` this will give us: ``` ../data/csv/file_utf-16.csv: Little-endian UTF-16 Unicode text ``` Another popular encodings are: - `cp1252` - `iso-8859-1` - `latin1` Python has option to check file encoding but it may be wrong in some cases like: ```python with open('../data/csv/file_utf-16.csv') as f: print(f) ``` result: ``` <_io.TextIOWrapper name='../data/csv/file_utf-16.csv' mode='r' encoding='UTF-8'> ``` ## 3: Solution of UnicodeDecodeError: skip encoding errors with encoding\_errors='ignore' Pandas `read_csv` has a parameter - `encoding_errors='ignore'` which defines how encoding errors are treated - to be skipped or raised. The parameter is described as: > How encoding errors are treated. Note: Important change in the new versions of Pandas: > Changed in version 1.3.0: encoding\_errors is a new argument. encoding has no longer an influence on how encoding errors are handled. Let's demonstrate how parameter of `read_csv` \- `encoding_errors` works: ```python from pathlib import Path import pandas as pd file = Path('../data/csv/file_utf-8.csv') file.write_bytes(b"\xe4\na\n1") # non utf-8 character df = pd.read_csv(file, encoding_errors='ignore') ``` The above will result into: | a | | - | | 1 | To prevent Pandas `read_csv` reading incorrect CSV data due to encoding use: `encoding_errors='strinct'` \- which is the default behavior: ```python df = pd.read_csv(file, encoding_errors='strict') ``` This will raise an error: > UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe4 in position 0: invalid continuation byte Another possible encoding error which can be raised by the same parameter is: > Pandas UnicodeEncodeError: 'charmap' codec can't encode character ## 4: Solution of UnicodeDecodeError: fix encoding errors with unicode\_escape The final solution to fix encoding errors like: - **UnicodeDecodeError** - **UnicodeEncodeError** is by using option `unicode_escape`. It can be described as: > Encoding suitable as the contents of a Unicode literal in ASCII-encoded Python source code, except that quotes are not escaped. Decode from Latin-1 source code. Beware that Python source code actually uses UTF-8 by default. Pandas `read_csv` and encoding can be used `'unicode_escape'` as: ```python df = pd.read_csv(file, encoding='unicode_escape') ``` to prevent encoding errors. ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/read%5Fcsv/pandas-read%5Fcsv-unicodedecodeerror-invalid-start-byte.ipynb?ref=datascientyst.com) - [pandas.read\_csv](https://pandas.pydata.org/pandas-docs/dev/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) - [BUG: read\_csv does not raise UnicodeDecodeError on non utf-8 characters](https://github.com/pandas-dev/pandas/issues/39450?ref=datascientyst.com) - [What’s new in 1.3.0 (July 2, 2021) - encoding\_errors](https://pandas.pydata.org/pandas-docs/stable/whatsnew/v1.3.0.html?ref=datascientyst.com) - [Text Encodings](https://docs.python.org/3/library/codecs.html?ref=datascientyst.com#text-encodings) ### How to Use set_index With MultiIndex Columns in Pandas URL: https://datascientyst.com/use-set_index-multiindex-columns-pandas/ Last updated: 2021-09-06T21:15:29.000Z Need to use the method `set_index` with MultiIndex columns in Pandas? If so you can find how to set single or multiple columns as index in Pandas DataFrame. To start let's create an example DataFrame with multi-level index for columns: ```python import pandas as pd cols = pd.MultiIndex.from_tuples([('company A', 'rank'), ('company A', 'points'), ('company B', 'rank'), ('company B', 'points')]) df = pd.DataFrame([[1,2,3,4], [2,3, 3,4]], columns=cols) ``` the DataFrame: | company A | company B | | | | --------- | --------- | ---- | ------ | | rank | points | rank | points | | 1 | 2 | 3 | 4 | | 2 | 3 | 3 | 4 | Before to show how to set the index for a given column let's check the column names by: ```python df.columns ``` the result is: ``` MultiIndex([('company A', 'rank'), ('company A', 'points'), ('company B', 'rank'), ('company B', 'points')], ) ``` So column name is a tuple of values: - `('company A', 'rank')` - `('company A', 'points')` Now in order to use the method `set_index` with columns we need to provide all levels from the hierarchical index. So **`set_index` applied on a single column**: ```python df.set_index([('company A', 'rank')]) ``` or getting the names with attribute `columns`: ```python df.set_index(df.columns[0]) ``` this will change the DataFrame to: | | company A | company B | | | ----------------- | --------- | --------- | ------ | | | points | rank | points | | (company A, rank) | | | | | 1 | 2 | 3 | 4 | | 2 | 3 | 3 | 4 | To use `set_index` with multiple columns in Pandas DataFrame we can apply next syntax: ```python df.set_index([('company A', 'rank'), ('company B', 'rank')]) ``` output: | | | company A | company B | | ----------------- | ----------------- | --------- | --------- | | | | points | points | | (company A, rank) | (company B, rank) | | | | 1 | 3 | 2 | 4 | | 2 | 3 | 3 | 4 | ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/multiindex/5.use-set%5Findex-multiindex-columns-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.set\_index](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.set%5Findex.html?ref=datascientyst.com) ### How to Flatten a MultiIndex in Pandas URL: https://datascientyst.com/flatten-multiindex-in-pandas/ Last updated: 2021-12-02T10:49:32.000Z Here are several approaches to **flatten hierarchical index in Pandas DataFrame**: (1) Flatten column MultiIndex with method `to_flat_index`: ```python df.columns = df.columns.to_flat_index() ``` (2) Flatten hierarchical index in DataFrame with `.get_level_values(0)`: ```python df.columns = df.columns.get_level_values(0) + '_' + df.columns.get_level_values(1) ``` (3) Pandas MultiIndex can be flatten with `reset_index(drop=True)`: ```python df.T.reset_index(drop=True).T ``` **MultiIndex can be flatten on rows and columns.** Next you can find several examples demonstrating how to use the above approaches. ## Step 1: Create DataFrame with MultiIndex Let's say that you have the following DataFrame with hierarchical index on columns: ```python import pandas as pd cols = pd.MultiIndex.from_tuples([('company', 'rank'), ('company', 'points')]) df = pd.DataFrame([[1,2], [3,4]], columns=cols) ``` data: | company | | | ------- | ------ | | rank | points | | 1 | 2 | | 3 | 4 | ## Step 2: Flatten column MultiIndex with method `to_flat_index` To flatten hierarchical index on columns or rows we can use the native Pandas method - `to_flat_index`. The method is described as: > Convert a MultiIndex to an Index of Tuples containing the level values. ```python df.columns = df.columns.to_flat_index() ``` This will change the MultiIndex to a normal index. From: ``` MultiIndex([('company', 'rank'), ('company', 'points')], ) ``` to: ``` Index([('company', 'rank'), ('company', 'points')], dtype='object') ``` ## Step 3: Flatten hierarchical index in DataFrame with `.get_level_values(0)` An alternative solution which gives control on the levels and the final format is - `.get_level_values(0)`. Instead of tuples we can get string concatenation of all levels of the MultiIndex. This method will return the levels accessed by their level number. So to access the highest level we can use 0, then 1 etc. The method is described as: > Return an Index of values for requested level. And to convert a MultiIndex levels to a simple index for columns we can combine all levels from the DataFrame: ```python df.columns = df.columns.get_level_values(0) + '_' + df.columns.get_level_values(1) ``` result: ``` Index(['company_rank', 'company_points'], dtype='object') ``` ## Step 4: Pandas flatten MultiIndex by `reset_index(drop=True)` Method **`reset_index` can flatten hierarchical index on rows and/or columns.** The usage for columns is a bit more complicated so we will share it as an example. The method will reset all levels and will reindex the columns ```python df.T.reset_index(drop=True).T ``` result: ``` RangeIndex(start=0, stop=2, step=1) ``` ## Step 5: Flatten MultiIndex in Pandas with list comprehension Finally let's cover the simple usage of Python list comprehension on column MultiIndex. This method is flexible and you have control of the output. Similar to Step 2 we are going to iterate through all levels and concatenate the values with symbol - `&`: ```python df.columns = [' & '.join(col).rstrip('_') for col in df.columns.values] ``` result: ``` Index(['company & rank', 'company & points'], dtype='object') ``` ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/multiindex/4.flatten-multiindex-in-pandas.ipynb?ref=datascientyst.com) - [pandas.MultiIndex.to\_flat\_index](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.MultiIndex.to%5Fflat%5Findex.html?ref=datascientyst.com) - [pandas.Index.get\_level\_values](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Index.get%5Flevel%5Fvalues.html?ref=datascientyst.com) ### How to Read Excel or CSV With Multiple Line Headers Using Pandas URL: https://datascientyst.com/read-excel-csv-multiple-line-headers-using-pandas/ Last updated: 2021-09-06T09:36:10.000Z In this quick Pandas tutorial, we'll cover how we can **read Excel sheet or CSV file with multiple header rowswith Python/Pandas.** Reading multi-line headers with Pandas creates a MultiIndex. Reading multiple headers from a CSV or Excel files can be done by using parameter - `header` of method `read_csv`: ```python import pandas as pd df = pd.read_csv('../data/csv/multine_header.csv', header=[0,1]) ``` In the rest of the article we will cover different examples and details about using `header=[0,1]`. ## Step 1: Sample CSV or Excel sheet To start lets create a simple CSV file named: `multine_header.csv` and show how we can **read the multi rows header with Pandas read\_csv method**. The content of the file is: ```python Date,Company A,Company A,Company B,Company A ,Rank,Points,Rank,Points 2021-09-06,1,7.9,2,6 2021-09-07,1,8.5,2,7 2021-09-08,2,8,1,8.1 ``` Data from the above file shown in a tabular form is(the same is if we read the CSV without the multi row header): | Date | Company A | Company A.1 | Company B | Company B.1 | | ---------- | --------- | ----------- | --------- | ----------- | | NaN | Rank | Points | Rank | Points | | 2021-09-06 | 1 | 7.9 | 2 | 6 | | 2021-09-07 | 1 | 8.5 | 2 | 7 | | 2021-09-08 | 2 | 8 | 1 | 8.1 | ## Step 2: Read CSV file with multiple headers To read the above CSV file which has two headers we can use `read_csv` with a combination of parameter `header`. The parameter is described as: > Row number(s) to use as the column names, and the start of the data. Default behavior is to infer the column names: if no names are passed the behavior is identical to header=0 and column names are inferred from the first line of the file, if column names are passed explicitly then the behavior is identical to header=None. So if a CSV file has two rows as a headers we can read them by: ```python import pandas as pd df = pd.read_csv('../data/csv/multine_header.csv', header=[0,1]) ``` Result: | Date | Company A | Company B | | | | -------------------- | --------- | --------- | ---- | ------ | | Unnamed: 0\_level\_1 | Rank | Points | Rank | Points | | 2021-09-06 | 1 | 7.9 | 2 | 6.0 | | 2021-09-07 | 1 | 8.5 | 2 | 7.0 | | 2021-09-08 | 2 | 8.0 | 1 | 8.1 | Now we can notice that DataFrame has two levels of columns. In next step we can find how to access the data. ### Read CSV file with 3 and more headers **To read CSV file with more than two rows as headers** we can use: ```python df = pd.read_csv('../data/csv/multine_header.csv', header=[0,1,2]) ``` ## Step 3: Access data from multi-line header DataFrame In order to access columns of the above DataFrame we need to use MultiIndex syntax. First let's find what are the column names by: ```python df.columns ``` The column names are presented as tuple pairs because they are a MultiIndex: ``` MultiIndex([( 'Date', 'Unnamed: 0_level_1'), ('Company A', 'Rank'), ('Company A', 'Points'), ('Company B', 'Rank'), ('Company B', 'Points')], ``` So in order to access column `Company A` \- `Rank` we will need to use the following syntax: ```python df[('Company A', 'Rank')] ``` result: ``` 0 1 1 1 2 2 Name: (Company A, Rank), dtype: int64 ``` and **to read multiple columns from multi-line header we can use**: ```python df[[('Company A', 'Rank'), ('Company B', 'Rank')]] ``` result: | Company A | Company B | | --------- | --------- | | Rank | Rank | | 1 | 2 | | 1 | 2 | | 2 | 1 | ## Step 4: Read Excel file with multiple headers Pandas can **read excel sheets with multiple headers** the same way as the CSV files. Below you can find the code for reading multiple headers from excel file: ```python pd.read_excel('../data/excel/multine_header.xlsx', sheet_name="multine_header", header=[0,1]) ``` Where the file name is: `multine_header.xlsx`, the sheet name is `multine_header`. To learn more about reading Excel files with Python and Pandas please check: [Read Excel XLS with Python Pandas](https://datascientyst.com/read-excel-xls-with-python-pandas/) ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/read%5Fcsv/read-excel-csv-multiple-line-headers-using-pandas.ipynb?ref=datascientyst.com) - [pandas.read\_csv](https://pandas.pydata.org/pandas-docs/dev/reference/api/pandas.read%5Fcsv.html?ref=datascientyst.com) - [Read Excel XLS with Python Pandas](https://datascientyst.com/read-excel-xls-with-python-pandas/) ### How to Highlight NaN Values in Pandas DataFrame URL: https://datascientyst.com/highlight-nan-values-pandas-dataframe/ Last updated: 2021-09-03T22:22:36.000Z Here are two ways to highlight `nan` values in a Pandas DataFrame: 1. highlight nan values in red - using `pd.isna` and `style.applymap` ```python df.style.applymap(lambda x: 'color: red' if pd.isna(x) else '') ``` 1. change background of nan values - comparing the value to itself ```python df.style.applymap(lambda x: '' if x==x else 'background-color: yellow') ``` Let's see several useful examples applying both ways in practice. ## Step 1: Create sample DataFrame Let's start with DataFrame with random numbers: ```python import pandas as pd import numpy as np df = pd.DataFrame(np.random.randint(0,100,size=(5, 5)), columns=list(range(0, 5))) ``` Data: | | 0 | 1 | 2 | 3 | 4 | | - | -- | -- | -- | -- | -- | | 0 | 48 | 12 | 94 | 73 | 13 | | 1 | 9 | 77 | 24 | 6 | 63 | | 2 | 63 | 38 | 49 | 86 | 39 | | 3 | 93 | 98 | 84 | 91 | 8 | | 4 | 59 | 55 | 10 | 64 | 87 | Let's set some of those cells to NaN values with generating pairs of coordinates: ```python import numpy as np randoms = np.random.choice(4,(5,2),replace=True) ``` result: ``` array([[0, 2], [3, 2], [2, 1], [1, 3], [1, 1]]) ``` set those coordinates to NaN values with simple loop and `df.loc`: ```python import numpy as np for pair in randoms: df.loc[pair[0],pair[1]] = np.nan ``` this will result into: | | 0 | 1 | 2 | 3 | 4 | | - | -- | ---- | ---- | ---- | -- | | 0 | 48 | 12.0 | NaN | 73.0 | 13 | | 1 | 9 | NaN | 24.0 | NaN | 63 | | 2 | 63 | NaN | 49.0 | 86.0 | 39 | | 3 | 93 | 98.0 | NaN | 91.0 | 8 | | 4 | 59 | 55.0 | 10.0 | 64.0 | 87 | ## Step 2: Highlight NaN values with lambda and pd.isna First lets color all NaN values in the DataFrame by using a lambda and `pd.isna`: ```python df.style.applymap(lambda x: 'color: red' if pd.isna(x) else '') ``` You can see the result below: | | 0 | 1 | 2 | 3 | 4 | | - | -- | --------- | --------- | --------- | -- | | 0 | 48 | 12.000000 | nan | 73.000000 | 13 | | 1 | 9 | nan | 24.000000 | nan | 63 | | 2 | 63 | nan | 49.000000 | 86.000000 | 39 | | 3 | 93 | 98.000000 | nan | 91.000000 | 8 | | 4 | 59 | 55.000000 | 10.000000 | 64.000000 | 87 | ## Step 3: Highlight NaN values by changing background In this step we are going to change the background of each NaN cell to yellow with `applymap`: ```python df.style.applymap(lambda x: '' if x==x else 'background-color: yellow') ``` result: | | 0 | 1 | 2 | 3 | 4 | | - | -- | --------- | --------- | --------- | -- | | 0 | 48 | 12.000000 | nan | 73.000000 | 13 | | 1 | 9 | nan | 24.000000 | nan | 63 | | 2 | 63 | nan | 49.000000 | 86.000000 | 39 | | 3 | 93 | 98.000000 | nan | 91.000000 | 8 | | 4 | 59 | 55.000000 | 10.000000 | 64.000000 | 87 | ## Step 4: Highlight NaN values in specific columns To change the color of NaN values only in selected columns we can use the parameter `subset` of method `style.applymap`. It can accept list of column names: ```python df.style.applymap(lambda x: '' if x==x else 'background-color: yellow', subset=[2,3]) ``` result: | | 0 | 1 | 2 | 3 | 4 | | - | -- | --------- | --------- | --------- | -- | | 0 | 48 | 12.000000 | nan | 73.000000 | 13 | | 1 | 9 | nan | 24.000000 | nan | 63 | | 2 | 63 | nan | 49.000000 | 86.000000 | 39 | | 3 | 93 | 98.000000 | nan | 91.000000 | 8 | | 4 | 59 | 55.000000 | 10.000000 | 64.000000 | 87 | ## Step 5: Highlight NaN values in specific columns and rows It's possible to select rows and columns in which NaN values to be highlighted. For this purpose we will use 2d input in order to select rows and columns: `subset=([0,1,2], slice(None))` ```python df.style.applymap(lambda x: '' if x==x else 'background-color: yellow', subset=([0,1,2], slice(None))) ``` result: | | 0 | 1 | 2 | 3 | 4 | | - | -- | --------- | --------- | --------- | -- | | 0 | 48 | 12.000000 | nan | 73.000000 | 13 | | 1 | 9 | nan | 24.000000 | nan | 63 | | 2 | 63 | nan | 49.000000 | 86.000000 | 39 | | 3 | 93 | 98.000000 | nan | 91.000000 | 8 | | 4 | 59 | 55.000000 | 10.000000 | 64.000000 | 87 | ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/nan/highlight-nan-values-pandas-dataframe.ipynb?ref=datascientyst.com) - [pandas.io.formats.style.Styler.applymap](https://pandas.pydata.org/docs/reference/api/pandas.io.formats.style.Styler.applymap.html?ref=datascientyst.com) - [pandas.isna](https://pandas.pydata.org/docs/reference/api/pandas.isna.html?highlight=isna&ref=datascientyst.com) ### How to Get most frequent values in Pandas DataFrame URL: https://datascientyst.com/get-most-frequent-values-pandas-dataframe/ Last updated: 2022-03-14T21:40:57.000Z In this short post, I'll show you how to **get most frequent values in Pandas DataFrame.** You can find also: [How to Get Top 10 Highest or Lowest Values in Pandas](https://datascientyst.com/get-top-10-highest-lowest-values-pandas/) To start, here are the ways **to get most frequent N values in your DataFrame**: ```python df['Magnitude'].value_counts() df['Magnitude'].mode() ``` In the next steps we will cover more details in simple examples ## Step 1: Create Sample DataFrame To start, let's create DataFrame with data from Kaggle:[Significant Earthquakes, 1965-2016](https://www.kaggle.com/usgs/earthquake-database?select=database.csv&ref=datascientyst.com). To learn more about Pandas and Kaggle: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) ```python import pandas as pd cols = ['Date', 'Time', 'Latitude', 'Longitude', 'Depth', 'Magnitude Type', 'Type', 'ID', 'Magnitude'] df = pd.read_csv(f'../data/earthquakes_1965_2016_database.csv.zip')[cols] ``` data: | Date | Latitude | Longitude | Depth | Magnitude Type | Magnitude | | ---------- | -------- | --------- | ----- | -------------- | --------- | | 01/02/1965 | 19.246 | 145.616 | 131.6 | MW | 6.0 | | 01/04/1965 | 1.863 | 127.352 | 80.0 | MW | 5.8 | | 01/05/1965 | \-20.579 | \-173.972 | 20.0 | MW | 6.2 | | 01/08/1965 | \-59.076 | \-23.557 | 15.0 | MW | 5.8 | | 01/09/1965 | 11.938 | 126.427 | 15.0 | MW | 5.8 | ## Step 2: Get Most Frequent value of Column in Pandas To get the most frequent value of a column we can use the method `mode`. It will return the value that appears most often. It can be multiple values. So to get the most frequent value in a single column - 'Magnitude' we can use: ```python df['Magnitude'].mode() ``` result: ``` 0 5.5 dtype: float64 ``` It has few parameters like: - `numeric_only` - `dropna` - `axis` If the **most frequent values are multiple they will be returned**: ```python df['Time'].mode() ``` result: ``` 0 02:56:58 1 14:09:03 dtype: object ``` ## Step 3: Get Most Frequent value for all columns in Pandas In order **to get the most frequent value for all columns we can iterate through all columns of the DataFrame**: ```python dfs = [] for col in df.columns: top_values = [] top_values = df[col].mode() dfs.append(pd.DataFrame({col: top_values}).reset_index(drop=True)) pd.concat(dfs, axis=1) ``` the most frequent values for all columns: | Date | Time | Depth | Magnitude Type | Type | Magnitude | | ---------- | -------- | ----- | -------------- | ---------- | --------- | | 03/11/2011 | 02:56:58 | 10.0 | MW | Earthquake | 5.5 | | NaN | 14:09:03 | NaN | NaN | NaN | NaN | Or get the most frequent values per dtype-s - for example only numeric columns: ```python from pandas.api.types import is_numeric_dtype dfs = [] for col in df.columns: top_values = [] if is_numeric_dtype(df[col]): top_values = df[col].mode() dfs.append(pd.DataFrame({col: top_values}).reset_index(drop=True)) pd.concat(dfs, axis=1) ``` The most frequent values for numeric columns | Depth | Magnitude | | ----- | --------- | | 10.0 | 5.5 | ## Step 4: Get N most frequent values in a column To get 5, 10 or N most frequent values in a single column we can use method `value_counts`: ```python df['Magnitude'].value_counts() ``` it will return the frequency for all values in the column: ``` 5.50 4685 5.60 3967 5.70 3079 5.80 2346 5.90 1947 6.00 1580 .... Name: Magnitude, dtype: int64 ``` To get the **N most frequent values only without the count** we can use: ```python n = 5 df['Magnitude'].value_counts().index.tolist()[:n] ``` result: ``` [5.5, 5.6, 5.7, 5.8, 5.9] ``` To get only the count: ```python n = 5 df['Magnitude'].value_counts().values.tolist()[:n] ``` result: \[4685, 3967, 3079, 2346, 1947\] To get the **single most frequent value from `value_counts` we can combine it with `idmax`**: ```python df['Magnitude'].value_counts().idxmax() ``` output: ``` 5.5 ``` **To get the count of the most frequent value:** ```python df['Magnitude'].value_counts().max() ``` output: ``` 4685 ``` ## Step 5: Get Top 10 most frequent values for all columns Finally let's cover the case - how to get the most frequent values for all columns in DataFrame. For this purpose we are going to use loop: ```python from pandas.api.types import is_categorical_dtype for col in df.columns: print(col, end=' - \n') print('_' * 50) if col in ['Magnitude'] or is_categorical_dtype(col): display(pd.DataFrame(df[col].astype('str').value_counts().sort_values(ascending=False).head(3))) else: display(pd.DataFrame(df[col].value_counts().sort_values(ascending=False).head(5))) ``` In the example above we are applying different behaviour for categorical columns. Because value\_counts can't be applied directly on them. result: Date - --- Date 03/11/2011 128 12/26/2004 51 02/27/2010 39 Time - --- Time 02:56:58 5 14:09:03 5 16:25:34 4 ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/0.guides/get-most-frequent-values-pandas-dataframe.ipynb?ref=datascientyst.com) - [pandas.DataFrame.mode](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.mode.html?ref=datascientyst.com) - [pandas.DataFrame.value\_counts](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.value%5Fcounts.html?ref=datascientyst.com) - [List of Aggregation Functions(aggfunc) for GroupBy in Pandas](https://datascientyst.com/list-aggregation-functions-aggfunc-groupby-pandas/) ### How to Check the Dtype of Column(s) in Pandas DataFrame URL: https://datascientyst.com/check-dtype-column-columns-pandas-dataframe/ Last updated: 2021-12-02T10:50:02.000Z To check the dtypes of single or multiple columns in Pandas you can use: ```python df.dtypes ``` Let's see other useful ways to check the dtypes in Pandas. ## Step 1: Create sample DataFrame To start, let's say that you have the date from earthquakes: | Date | Time | Depth | Magnitude Type | Type | Magnitude | Depth\_int | | ------------------------- | -------- | ----- | -------------- | ---------- | --------- | ---------- | | 1965-01-02 00:00:00+00:00 | 13:44:18 | 131.6 | MW | Earthquake | 6.0 | 131 | | 1965-01-04 00:00:00+00:00 | 11:29:49 | 80.0 | MW | Earthquake | 5.8 | 80 | | 1965-01-05 00:00:00+00:00 | 18:05:58 | 20.0 | MW | Earthquake | 6.2 | 20 | | 1965-01-08 00:00:00+00:00 | 18:49:43 | 15.0 | MW | Earthquake | 5.8 | 15 | | 1965-01-09 00:00:00+00:00 | 13:32:50 | 15.0 | MW | Earthquake | 5.8 | 15 | Data is available from Kaggle: [Significant Earthquakes, 1965-2016](https://www.kaggle.com/usgs/earthquake-database?select=database.csv&ref=datascientyst.com). How to read and convert Kaggle data to Pandas DataFrame: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) ## Step 2: Get dtypes for all columns in DataFrame To get dtypes details for the whole DataFrame you can use attribute - `dtypes`: ```python df.dtypes ``` the result is: ``` Date datetime64[ns, UTC] Time object Depth float64 Magnitude Type object Type object Magnitude float64 Depth_int int64 dtype: object ``` we can see several different types like: - `datetime64[ns, UTC]` \- it's used for dates; explicit conversion may be needed in some cases - `float64` / `int64` \- numeric data - `object` \- strings and other ## Step 3: Short explanation of dtypes in Pandas Let's briefly cover some dtypes and their usage with simple examples. **Table of the most used dtypes in Pandas:** | Pandas dtype | Data Type | Description | Example | Creation | | ------------ | --------- | ------------------------------------- | ------------------------------- | ------------------------------------ | | bool | bool | Boolean values – True or False | True | pd.BooleanDtype() | | category | NA | Limited list of values (can be fixed) | \[‘red’, ‘blue’\] | pd.Categorical(\[1, 2, 3, 1, 2, 3\]) | | datetime64 | datetime | Datetime (conversion is needed) | 2020-11-16 22:50:18.092888+0000 | to\_datetime(df\['date'\]) | | float64 | float | Floating point numbers | 80.5 | df.astype('float64') | | int64 | int | Integer numbers | 8 | df.astype('int64') | | object | strings | String, text and other | Red Pandas | | | timedelta | timedelta | Duration between two dates or times | 0 days 00:00:00.000000001 | pd.Timedelta(42, unit='ns') | More information about them can be found on this link: [Pandas User Guide dtypes](https://pandas.pydata.org/docs/user%5Fguide/basics.html?ref=datascientyst.com#dtypes). Pandas offers a wide range of features and methods in order to read, parse and convert between different dtypes. The most popular conversion methods are: - `to_datetime(df['date'])` - `to_timedelta(df['timdelta'])` - `to_numeric(df['amount'])` - `df['amount'].astype('int32')` ## Step 4: Check if column is numeric, datetime, categorical etc In this step we are going to see how we can check if a given column is numerical or categorical. For this purpose Pandas offers a bunch of methods like: - `is_string_dtype` - `is_dict_like` - `is_list_like` - `is_numeric_dtype` - `is_datetime64_dtype` To find all methods you can check the official Pandas docs: [pandas.api.types.is\_datetime64\_any\_dtype](https://pandas.pydata.org/docs/reference/api/pandas.api.types.is%5Fdatetime64%5Fany%5Fdtype.html?ref=datascientyst.com) To check if a column has numeric or datetime dtype we can: ```python from pandas.api.types import is_numeric_dtype is_numeric_dtype(df['Depth_int']) ``` result: ``` True ``` for datetime exists several options like: `is_datetime64_ns_dtype` or `is_datetime64_any_dtype`: ```python from pandas.api.types import is_datetime64_any_dtype is_datetime64_any_dtype(df['Date']) ``` result: ``` True ``` ## Step 5: List all numeric/datetime columns in Pandas DataFrame If you like to list only numeric/datetime or other type of columns in a DataFrame you can use method `select_dtypes`: **including** ```python df.select_dtypes(include=['float64']).columns ``` result of the operation: ``` Index(['Depth', 'Magnitude'], dtype='object') ``` excluding columns by dtype: ```python df.select_dtypes(exclude=['float64','datetime']).columns ``` result: ``` Index(['Date', 'Time', 'Magnitude Type', 'Type', 'Depth_int'], dtype='object') ``` ## Step 6: Filter columns by dtype and name in Pandas DataFrame As an alternative solution you can construct a loop over all columns. Then you can check the dtype and the name of the column. Below we are listing all numeric column which name has word 'Depth': ```python from pandas.api.types import is_numeric_dtype for col in df.columns: if is_numeric_dtype(df[col]) and 'Depth' in col: print(col) ``` As a result you will get a list of all numeric columns: ``` Depth Depth_int ``` Instead of printing their names you can do something. ## Step 7: Apply function on numeric columns only To apply function to numeric or datetime columns only you can use the method `select_dtypes` in combination with `apply`. The function below will iterate over all numeric columns and double the value: ```python def double_n(x): return x * df.select_dtypes(include=['float64']).apply(double_n) ``` ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/column/3.check-dtype-column-columns-pandas-dataframe.ipynb?ref=datascientyst.com) - [pandas.DataFrame.dtypes](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.dtypes.html?ref=datascientyst.com) - [dtypes](https://pandas.pydata.org/docs/user%5Fguide/basics.html?ref=datascientyst.com#dtypes) - [pandas.DataFrame.astype](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.astype.html?ref=datascientyst.com) - [pandas.api.types.is\_datetime64\_any\_dtype](https://pandas.pydata.org/docs/reference/api/pandas.api.types.is%5Fdatetime64%5Fany%5Fdtype.html?ref=datascientyst.com) - [pandas.DataFrame.select](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.select%5Fdtypes.html?ref=datascientyst.com) ### How to Get Top 10 Highest or Lowest Values in Pandas URL: https://datascientyst.com/get-top-10-highest-lowest-values-pandas/ Last updated: 2021-12-02T10:50:37.000Z In this short guide, I'll show you how to **get top 5, 10 or N values in Pandas DataFrame.** You can find also how to **print top/bottom values for all columns** in a DataFrame. If you need the most frequent values in a column or DataFrame you can check: [How to Get Most Frequent Values in Pandas Dataframe](https://datascientyst.com/p/a113714c-59ac-4be5-9375-cf6c5c0479ac/) To start, here is the syntax that you may apply in order to **get top 10 biggest numeric values in your DataFrame**: ```python df['Magnitude'].nlargest(n=10) ``` To list the **top 10 lowest values in DataFrame you can use:** ```python df.nlargest(n=5, columns=['Magnitude', 'Depth']) ``` In the next section, I’ll show you more examples and other techniques in order **to get top/bottom values in DataFrame**. ## Step 1: Create Sample DataFrame In this example we are going to use real data available from Kaggle:[Significant Earthquakes, 1965-2016](https://www.kaggle.com/usgs/earthquake-database?select=database.csv&ref=datascientyst.com). To learn more about reading Kaggle data with Python and Pandas: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) ```python import pandas as pd cols = ['Date', 'Time', 'Latitude', 'Longitude', 'Depth', 'Magnitude Type', 'Type', 'ID', 'Magnitude'] df = pd.read_csv(f'../data/earthquakes_1965_2016_database.csv.zip')[cols] ``` data: | Date | Latitude | Longitude | Depth | Magnitude Type | Magnitude | | ---------- | -------- | --------- | ----- | -------------- | --------- | | 01/02/1965 | 19.246 | 145.616 | 131.6 | MW | 6.0 | | 01/04/1965 | 1.863 | 127.352 | 80.0 | MW | 5.8 | | 01/05/1965 | \-20.579 | \-173.972 | 20.0 | MW | 6.2 | | 01/08/1965 | \-59.076 | \-23.557 | 15.0 | MW | 5.8 | | 01/09/1965 | 11.938 | 126.427 | 15.0 | MW | 5.8 | ## Step 2: Get Top 10 biggest/lowest values for single column You can use functions: - `nsmallest` \- return the first n rows ordered by columns in ascending order - `nlargest` \- return the first n rows ordered by columns in descending order to get the top N highest or lowest values in Pandas DataFrame. So let's say that we would like to find the strongest earthquakes by magnitude. From the data above we can use the following syntax: ```python df['Magnitude'].nlargest(n=10) ``` the result is: ``` 17083 9.1 20501 9.1 19928 8.8 16 8.7 17329 8.6 21219 8.6 ... Name: Magnitude, dtype: float64 ``` How does it return the top 10 values? It orders all values in ascending or descending order. Then will return 5, 10 or N values from this order. ## Step 3: Get Top 10 biggest/lowest values - duplicates As you can see in Step 2 there are duplicates in the results. `nsmallest` / `nlargest` have a parameter: - `keep{‘first’, ‘last’, ‘all’}, default ‘first’` which helps to deal with duplicates. Let's say that you would like to get 5 smallest values by `Depth` of the earthquake: ```python df.nsmallest(n=5, columns=['Depth']) ``` result: | Date | Time | Depth | Magnitude Type | Type | ID | Magnitude | | ---------- | -------- | ------- | -------------- | ----------------- | ---------- | --------- | | 06/28/1992 | 12:00:45 | \-1.100 | ML | Earthquake | CI3043549 | 5.77 | | 06/28/1992 | 11:57:34 | \-0.097 | MW | Earthquake | CI3031111 | 7.30 | | 07/21/1986 | 22:07:16 | \-0.076 | ML | Earthquake | NC95459 | 5.60 | | 02/16/1973 | 05:02:58 | 0.000 | MB | Explosion | USP00000JC | 5.60 | | 07/23/1973 | 01:22:58 | 0.000 | MB | Nuclear Explosion | USP00002TP | 6.30 | It seems that we have many records with Depth 0\. In order to get all records which have Depth 0 we can use: ```python df.nsmallest(n=5, columns=['Depth'], keep='all') ``` the result will return the previous ones plus the duplicated: ``` 173 rows × 7 columns ``` While if we use `keep='last'` we will get again 5 rows as a result. ## Step 4: Get Top N values in multiple columns You can get top 5, 10 or N records from multiple columns with: - `nsmallest` - `nlargest` To do so you can use the following syntax: ```python df.nlargest(n=5, columns=['Magnitude', 'Depth']) ``` result: | Date | Time | Depth | Magnitude Type | Type | Magnitude | | ---------- | -------- | ----- | -------------- | ---------- | --------- | | 12/26/2004 | 00:58:53 | 30.0 | MW | Earthquake | 9.1 | | 03/11/2011 | 05:46:24 | 29.0 | MWW | Earthquake | 9.1 | | 02/27/2010 | 06:34:12 | 22.9 | MWW | Earthquake | 8.8 | | 02/04/1965 | 05:01:22 | 30.3 | MW | Earthquake | 8.7 | | 03/28/2005 | 16:09:37 | 30.0 | MWW | Earthquake | 8.6 | ## Step 5: How do `nsmallest` and `nlargest` work So to summarise the algorithm: - order all values - ascending or descending - find top N values - filter out duplicates ( depends on parameter `keep` ) - return N top values Both methods will work only with numeric types. If you try to use them on columns with dtype object you will get an error: > TypeError: Column 'Magnitude Type' has dtype object, cannot use method 'nlargest' with this dtype ## Step 6: Display Top values for all numeric columns in DataFrame If you like to get top values for all columns you can use the next syntax: ```python from pandas.api.types import is_numeric_dtype dfs = [] for col in df.columns: top_values = [] if is_numeric_dtype(df[col]): top_values = df[col].nlargest(n=7) dfs.append(pd.DataFrame({col: top_values}).reset_index(drop=True)) pd.concat(dfs, axis=1) ``` This will return the top values per column as a new DataFrame: | Depth | Magnitude | | ----- | --------- | | 700.0 | 9.1 | | 691.6 | 9.1 | | 690.0 | 8.8 | | 688.0 | 8.7 | | 687.6 | 8.6 | ## Step 7: `nsmallest` and `nlargest` and datetime It's possible to use methods **`nsmallest` and `nlargest` with datetime**. You need to be sure that the target column is from type: `datetime64[ns, UTC]` If you need to convert string to datetime column you can use: ```python df['Date'] = pd.to_datetime(df['Date'], utc=True) ``` and then use the `nlargest`: ```python df['Date'].nlargest() ``` ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/0.guides/1.get-top-10-highest-lowest-values-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.nlargest](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.nlargest.html?ref=datascientyst.com) - [pandas.DataFrame.nsmallest](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.nsmallest.html?ref=datascientyst.com) ### How to Replace Values in Column Based On Another DataFrame in Pandas URL: https://datascientyst.com/replace-values-column-based-another-dataframe-pandas/ Last updated: 2026-01-23T17:18:27.000Z In this quick tutorial, we'll cover how we can **replace values in a column based on values from another DataFrame in Pandas**. We can use the following syntax to margin on a single axis column or row in Pandas: **(1) matching indices** ```python df2.loc[:, ['ID']] = df1[['ID']] ``` **(2) non matching indices** ```python df3.loc[:, ['Latitude', 'Longitude']] = df1[['Latitude', 'Longitude']] ``` Mapping the values from another DataFrame, depends on several factors like: - Index matching - Update only NaN values, add new column or replace everything In this article, we are going to answer on all questions in a different steps. ## Step 1: Create sample DataFrame For this article we are going to use data from Kaggle: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) Reading the initial data: ```python import pandas as pd df1 = pd.read_csv(f'../data/earthquakes_1965_2016_database.csv.zip') ``` Next we are going to create a second DataFrame. We are going to replace the values from the first to the second one. The second DataFrame will be created from the values of the first for one - `Latitude` and `Longitude` and returning the country. ```python import geocoder def geo_rev(x): g = geocoder.osm([x.Latitude, x.Longitude], method='reverse').json if g: return g.get('country') else: return 'no country' df2 = pd.DataFrame({'country': df1[['Latitude', 'Longitude']].apply(geo_rev, axis=1)}) ``` So far we have two DataFrames: - `df1` | | Date | Latitude | Longitude | Depth | ID | | ----- | ---------- | -------- | ---------- | ----- | ---------- | | 23405 | 12/27/2016 | 45.7192 | 26.5230 | 97.0 | US10007N3R | | 23406 | 12/28/2016 | 38.3754 | \-118.8977 | 10.8 | NN00570709 | | 23407 | 12/28/2016 | 38.3917 | \-118.8941 | 12.3 | NN00570710 | | 23408 | 12/28/2016 | 38.3777 | \-118.8957 | 8.8 | NN00570744 | | 23409 | 12/28/2016 | 36.9179 | 140.4262 | 10.0 | US10007NAF | - `df2` | | country | | ----- | ------------- | | 23405 | România | | 23406 | United States | | 23407 | United States | | 23408 | United States | | 23409 | 日本 | ## Step 2: Replace Values with matching indices To **add a single column - `ID` we can use method `loc` in order to set the values from DataFrame `df1` to `df2`**: ```python df2.loc[:, ['ID']] = df1[['ID']] ``` This is possible because **both DataFrames have identical indices and shapes**. Otherwise error or unexpected results might happen. If you need to **replace values for multiple columns from another DataFrame - this is the syntax**: ```python df2.loc[:, ['Latitude', 'Longitude']] = df1[['Latitude', 'Longitude']] ``` The two columns are added from `df1` to `df2`: | | country | ID | Latitude | Longitude | | ----- | ------------- | ---------- | -------- | ---------- | | 23405 | România | US10007N3R | 45.7192 | 26.5230 | | 23406 | United States | NN00570709 | 38.3754 | \-118.8977 | | 23407 | United States | NN00570710 | 38.3917 | \-118.8941 | | 23408 | United States | NN00570744 | 38.3777 | \-118.8957 | | 23409 | 日本 | US10007NAF | 36.9179 | 140.4262 | ## Step 3: Replace Values with non matching indices What will happen if the **indexes do not match**? Let's create one more DataFrame `df3` which is a copy of 2. For this new DataFrame we are going to reset the index by: ```python df3 = df2.head(7).reset_index().copy() ``` If we try to use the same technique for setting values from another DataFrame we will get only `NaN` values: ```python df3.loc[:, ['Latitude', 'Longitude']] = df1[['Latitude', 'Longitude']] df3 ``` because the indices doesn't match: | | index | country | ID | Latitude | Longitude | | - | ----- | ------------- | ---------- | -------- | --------- | | 0 | 23405 | România | US10007N3R | NaN | NaN | | 1 | 23406 | United States | NN00570709 | NaN | NaN | | 2 | 23407 | United States | NN00570710 | NaN | NaN | | 3 | 23408 | United States | NN00570744 | NaN | NaN | | 4 | 23409 | 日本 | US10007NAF | NaN | NaN | In order to make it work we need to modify the code. We are going to use column `ID` as a reference between the two DataFrames. Two columns `'Latitude', 'Longitude'` will be set from DataFrame `df1` to `df2`. So to **replace values from another DataFrame when different indices** we can use: ```python col = 'ID' cols_to_replace = ['Latitude', 'Longitude'] df3.loc[df3[col].isin(df1[col]), cols_to_replace] = df1.loc[df1[col].isin(df3[col]),cols_to_replace].values ``` Now the values are correctly set: | | country | ID | Latitude | Longitude | | ----- | ------------- | ---------- | -------- | ---------- | | 23405 | România | US10007N3R | 45.7192 | 26.5230 | | 23406 | United States | NN00570709 | 38.3754 | \-118.8977 | | 23407 | United States | NN00570710 | 38.3917 | \-118.8941 | | 23408 | United States | NN00570744 | 38.3777 | \-118.8957 | | 23409 | 日本 | US10007NAF | 36.9179 | 140.4262 | ## Step 4: Insert new column with values from another DataFrame by merge You can use **Pandas `merge` function in order to get values and columns from another DataFrame.** For this purpose you will need to have reference column between both DataFrames or use the index. In this example we are going to use reference column `ID` \- we will merge `df1` left join on `df4`. It's important to mention two points: - `ID` \- should be unique value - the column which is updated should not exists in the first DataFrame - `df4` ```python df4 = df4.merge(df1,on='ID',how="left") ``` the result is: | country | ID | Date | Time | Latitude | Longitude | Depth | | ------------- | ---------- | ---------- | -------- | -------- | ---------- | ----- | | România | US10007N3R | 12/27/2016 | 23:20:56 | 45.7192 | 26.5230 | 97.0 | | United States | NN00570709 | 12/28/2016 | 08:18:01 | 38.3754 | \-118.8977 | 10.8 | So all columns which has a match on the `ID` \- this column is unique per row - will get corresponding data from `df1` ## Step 5: Update missing Values from Another DataFrame In this step we are going to **update only missing values in a column in one DataFrame from another**. So let's have next DataFrame: ```python import numpy as np df5.loc[:, ['ID', 'Longitude']] = df1[['ID', 'Longitude']] df5.iloc[[1,3,4], -1] = np.NaN df5.head(5) ``` with few missing values in column `Longitude`: | | country | ID | Longitude | | ----- | ------------- | ---------- | ---------- | | 23405 | România | US10007N3R | 26.5230 | | 23406 | United States | NN00570709 | NaN | | 23407 | United States | NN00570710 | \-118.8941 | | 23408 | United States | NN00570744 | NaN | | 23409 | 日本 | US10007NAF | NaN | To update the values in `df5` from `df1` we can merge both DataFrames on a reference column - `ID` (should be unique). Then we are going to fill all missing values with the values from the column `Longitude_x` of `df1` to the new column of `df5` \- `Longitude_y`. Next we are going to drop the column from `df1` and rename the new column: ```python df5 = df5.merge(df1[['Longitude', 'ID']],on='ID',how="left") df5['Longitude_y'] = df5['Longitude_y'].fillna(df5['Longitude_x']) df5.drop(["Longitude_x"], inplace=True, axis=1) df5.rename(columns={'Longitude_y':'Longitude'},inplace=True) ``` result: | | country | ID | Longitude | | - | ------------- | ---------- | ---------- | | 0 | România | US10007N3R | 26.5230 | | 1 | United States | NN00570709 | \-118.8977 | | 2 | United States | NN00570710 | \-118.8941 | | 3 | United States | NN00570744 | \-118.8957 | | 4 | 日本 | US10007NAF | 140.4262 | ## Step 6: Update column from another column with np.where Finally let's cover how column can be **added or updated from another column or DataFrame with `np.where`.** First we will check if one column contains a value and create another column: ```python import numpy as np df1['new_lon'] = np.where(df1['Longitude']>10,df1['Longitude'],np.nan) ``` If we have two DataFrames we can use similar syntax as follow: ```python df1['name'] = np.where(df2['Longitude']==1,df2['name'],df1['name']) ``` What is going on here is the following: - check if DataFrame `df2` contains rows with value 1 - all the rest will be taken from `df1` - create new column in `df1` with the result from previous `np.where` Note that the two DataFrames should have the same number of rows. ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/replace/1.replace-values-column-based-another-dataframe-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.loc](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.loc.html?ref=datascientyst.com) - [pandas.DataFrame.merge](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.merge.html?ref=datascientyst.com) ### List of Aggregation Functions(aggfunc) for GroupBy in Pandas URL: https://datascientyst.com/list-aggregation-functions-aggfunc-groupby-pandas/ Last updated: 2023-04-29T06:43:57.000Z In this article, you can find the list of the **available aggregation functions for groupby in Pandas**: - `count` / `nunique` – non-null values / count number of unique values - `min` / `max` – minimum/maximum - `first` / `last` \- return first or last value per group - `unique` \- all unique values from the group - `std` – standard deviation - `sum` – sum of values - `mean` / `median` / `mode` – mean/median/mode - `var` \- unbiased variance - `mad` \- mean absolute deviation - `skew` \- unbiased skew - `sem` \- standard error of the mean - `quantile` Those functions can be used with `groupby` in order to return statistical information about the groups. In the next section we will cover all aggregation functions with simple examples. ![pandas-aggfunc-stats-functions-min-max-group](https://datascientyst.com/content/images/2022/10/pandas-aggfunc-stats-functions-min-max-group.opti.webp) ## Step 1: Create DataFrame for aggfunc Let us use the earthquake dataset. We are going to create new column `year_month` and `groupby` by it: ```python import pandas as pd df = pd.read_csv(f'../data/earthquakes_1965_2016_database.csv.zip') cols = ['Date', 'Time', 'Latitude', 'Longitude', 'Depth', 'Magnitude Type', 'Type', 'ID'] df = df[cols] ``` result: | Date | Latitude | Longitude | Depth | Type | | ---------- | -------- | --------- | ----- | ---------- | | 01/02/1965 | 19.246 | 145.616 | 131.6 | Earthquake | | 01/04/1965 | 1.863 | 127.352 | 80.0 | Earthquake | | 01/05/1965 | \-20.579 | \-173.972 | 20.0 | Earthquake | | 01/08/1965 | \-59.076 | \-23.557 | 15.0 | Earthquake | | 01/09/1965 | 11.938 | 126.427 | 15.0 | Earthquake | Next we are going to create new column with information - combination of the year and the month: ```python df['Date'] = pd.to_datetime(df['Date'], utc=True) df['year_month'] = df['Date'].dt.to_period('M') ``` ## Step 2: Pandas describe DataFrame In each step we will see examples of using each of the aggregating functions associated with Pandas groupby function. We will start with the method `describe`. This method returns basic information about the column: ```python df.groupby('year_month')['Depth'].describe() ``` `describe` returns multiple aggfunc-s like: count, mean, std, min, max: | count | mean | std | min | 25% | 50% | 75% | max | | ----- | ---------- | ---------- | ---- | ------ | ---- | ------- | ----- | | 13.0 | 101.115385 | 152.237697 | 15.0 | 20.000 | 35.0 | 95.000 | 565.0 | | 54.0 | 47.712963 | 80.122216 | 10.0 | 20.075 | 25.1 | 34.375 | 482.9 | | 38.0 | 62.055263 | 97.890288 | 10.0 | 25.000 | 31.3 | 47.850 | 560.8 | | 33.0 | 112.163636 | 183.812083 | 10.0 | 25.000 | 30.7 | 64.700 | 635.0 | | 22.0 | 80.972727 | 114.894839 | 10.0 | 25.875 | 35.0 | 105.000 | 553.8 | ![](https://datascientyst.com/content/images/2021/08/list-aggregation-functions-aggfunc-groupby-pandas.png) ## Step 3: Pandas all aggfunc for DataFrame In this step you can find examples for all aggfunc-s applied on a DataFrame. The list of the functions is below. Note that by default method `groupby` will exclude all `NaN` values. In order to change this behavior you can use parameter - `dropna=False` ```python aggfuncs = [ 'count', 'sum', 'sem', 'skew', 'mean', 'min', 'max', 'std', 'quantile', 'nunique', 'mad', 'size', pd.Series.mode, 'var', 'unique'] df.groupby('year_month', dropna=False)['Depth'].agg(aggfuncs) ``` result: | year\_month | 1965-01 | 1965-02 | | ----------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | count | 13 | 54 | | sum | 1314.5 | 2576.5 | | sem | 42.22314 | 10.903253 | | skew | 2.744523 | 4.256526 | | mean | 101.115385 | 47.712963 | | min | 15.0 | 10.0 | | max | 565.0 | 482.9 | | std | 152.237697 | 80.122216 | | quantile | 35.0 | 25.1 | | nunique | 9 | 30 | | mad | 95.56213 | 39.163306 | | size | 13 | 54 | | mode | 20.0 | 25.0 | | var | 23176.31641 | 6419.569451 | | unique | \[131.6, 80.0, 20.0, 15.0, 35.0, 95.0, 565.0, 227.9, 55.0\] | \[482.9, 15.0, 10.0, 30.3, 30.0, 25.0, 20.0, 24.0, 31.8, 39.5, 30.4, 17.8, 27.7, 30.1, 37.4, 17.5, 22.5, 25.2, 17.7, 32.5, 200.0, 100.0, 53.5, 340.0, 55.0, 35.0, 40.4, 40.0, 20.3, 160.0\] | ## Step 4: Pandas aggfunc - Count, Nunique, Size, Unique In this step we will cover 4 aggregation functions: - `count` \- compute count of group, excluding missing values - `size` \- compute group sizes - `unique` \- return unique values - `nunique` \- return number of unique elements in the group. Example of using the functions and the result: ```python aggfuncs = [ 'count', 'size', 'nunique', 'unique'] df.groupby('year_month')['Depth'].agg(aggfuncs) ``` output: | | count | size | nunique | unique | | ----------- | ----- | ---- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | year\_month | | | | | | 1965-01 | 13 | 13 | 9 | \[131.6, 80.0, 20.0, 15.0, 35.0, 95.0, 565.0, 227.9, 55.0\] | | 1965-02 | 54 | 54 | 30 | \[482.9, 15.0, 10.0, 30.3, 30.0, 25.0, 20.0, 24.0, 31.8, 39.5, 30.4, 17.8, 27.7, 30.1, 37.4, 17.5, 22.5, 25.2, 17.7, 32.5, 200.0, 100.0, 53.5, 340.0, 55.0, 35.0, 40.4, 40.0, 20.3, 160.0\] | | 1965-03 | 38 | 38 | 24 | \[30.0, 40.0, 33.6, 105.2, 10.0, 15.0, 14.8, 25.7, 200.0, 35.0, 560.8, 45.0, 50.0, 20.0, 207.8, 32.1, 25.0, 224.9, 28.9, 30.5, 48.8, 55.0, 70.0, 75.0\] | | 1965-04 | 33 | 33 | 22 | \[60.0, 10.0, 30.7, 25.0, 20.0, 39.2, 50.0, 65.0, 17.5, 543.7, 635.0, 570.0, 421.7, 165.0, 15.0, 35.0, 70.0, 30.0, 36.8, 37.1, 64.7, 480.0\] | | 1965-05 | 22 | 22 | 16 | \[15.0, 22.5, 90.0, 110.0, 10.0, 50.0, 153.1, 25.0, 35.0, 68.4, 120.0, 553.8, 28.5, 30.1, 125.0, 150.0\] | ## Step 5: Pandas aggfunc - First and Last There are two functions which can return the first or the last value of the group. They are: - `first` \- compute first of group values - `last` \- compute first of group values Example of their usage: ```python aggfuncs = [ 'first', 'last'] df.groupby('year_month')['Depth'].agg(aggfuncs) ``` result: | | first | last | | ----------- | ----- | ----- | | year\_month | | | | 1965-01 | 131.6 | 55.0 | | 1965-02 | 482.9 | 10.0 | | 1965-03 | 30.0 | 75.0 | | 1965-04 | 60.0 | 480.0 | | 1965-05 | 15.0 | 150.0 | ## Step 6: Pandas aggfunc - Sum, Min, Max For numeric or datetime columns we can get the minimum, maximum or the sum by those aggfunc-s: - `sum` \- compute sum of group values - `min` \- compute min of group values - `max` \- compute max of group values How to get the sum, maximum and the minimum per group: ```python aggfuncs = [ 'sum', 'min', 'max'] df.groupby('year_month')['Depth'].agg(aggfuncs) ``` the result is: | | sum | min | max | | ----------- | ------ | ---- | ----- | | year\_month | | | | | 1965-01 | 1314.5 | 15.0 | 565.0 | | 1965-02 | 2576.5 | 10.0 | 482.9 | | 1965-03 | 2358.1 | 10.0 | 560.8 | | 1965-04 | 3701.4 | 10.0 | 635.0 | | 1965-05 | 1781.4 | 10.0 | 553.8 | ## Step 7: Pandas aggfunc - Mean, Median, Mode There are several very important statistics which are: - The mean is the average of a group values - The mode is the most common number in a group - The median is the middle of the group values They are implemented in Pandas as functions: - `mean` \- compute mean of groups, excluding missing values - `pd.Series.mode` \- return the mode(s) of the Series. - `median` \- compute median of groups, excluding missing values. They can be compute on Pandas `groupby` object by next syntax: ```python aggfuncs = [ 'mean', 'median', pd.Series.mode] df.groupby('year_month')['Depth'].agg(aggfuncs) ``` result: | | mean | median | mode | | ----------- | ---------- | ------ | ---- | | year\_month | | | | | 1965-01 | 101.115385 | 35.0 | 20.0 | | 1965-02 | 47.712963 | 25.1 | 25.0 | | 1965-03 | 62.055263 | 31.3 | 30.0 | | 1965-04 | 112.163636 | 30.7 | 25.0 | | 1965-05 | 80.972727 | 35.0 | 35.0 | ## Step 7: Pandas aggfunc - STD, MAD, Var Another important methods in statistics are: - `std` \- compute **standard deviation of groups**, excluding missing value - `var` \- compute **variance of groups**, excluding missing values - `mad` \- return the **mean absolute deviation** of the values over the requested axis How to calculate the standard deviation, variance and mean absolute deviation of groups: ```python aggfuncs = ['mad', 'std', 'var'] df.groupby('year_month')['Depth'].agg(aggfuncs) ``` output: | | mad | std | var | | ----------- | ---------- | ---------- | ------------ | | year\_month | | | | | 1965-01 | 95.562130 | 152.237697 | 23176.316410 | | 1965-02 | 39.163306 | 80.122216 | 6419.569451 | | 1965-03 | 53.121745 | 97.890288 | 9582.508485 | | 1965-04 | 129.843526 | 183.812083 | 33786.881761 | | 1965-05 | 66.826446 | 114.894839 | 13200.823983 | ## Step 7: Pandas aggfunc - Skew, Sem, quantile Let's check few other functions which are not very popular like: - `skew` \- return unbiased skew over requested axis - `sem` \- compute standard error of the mean of groups, excluding missing values - `quantile` \- return group values at the given quantile, a la numpy.percentile ```python aggfuncs = [ 'skew', 'sem', 'quantile'] df.groupby('year_month')['Depth'].agg(aggfuncs) ``` result: | | skew | sem | quantile | | ----------- | -------- | --------- | -------- | | year\_month | | | | | 1965-01 | 2.744523 | 42.223140 | 35.0 | | 1965-02 | 4.256526 | 10.903253 | 25.1 | | 1965-03 | 4.036620 | 15.879902 | 31.3 | | 1965-04 | 2.054374 | 31.997576 | 30.7 | | 1965-05 | 3.604639 | 24.495662 | 35.0 | ## Step 8: User defined aggfunc for groupby It's possible in Pandas to define your own aggfunc and use it with a groupby method. In the next example we will define a function which will compute the NaN values in each group: ```python def countna(x): return (x.isna()).sum() df.groupby('year_month')['Depth'].agg([countna]) ``` result: | | countna | | ----------- | ------- | | year\_month | | | 1965-01 | 0 | | 1965-02 | 0 | | 1965-03 | 0 | | 1965-04 | 0 | | 1965-05 | 0 | ## Step 9: Pandas aggfuncs from scipy or numpy Finally let's check how to use aggregation functions with `groupby` from `scipy` or `numpy` Below you can find a `scipy` example applied on Pandas `groupby` object: ```python from scipy import stats df.groupby('year_month')['Depth'].agg(lambda x: stats.mode(x)[0]) ``` result: ``` year_month 1965-01 20.0 1965-02 25.0 1965-03 30.0 1965-04 25.0 ``` Example for `numpy.count_nonzero` method used with Pandas `groupby` method: ```python import numpy as np df.groupby('year_month')['Depth'].agg(np.count_nonzero) ``` Output: ``` year_month 1965-01 13 1965-02 54 1965-03 38 1965-04 33 ``` ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/aggfunc/1.list-aggregation-functions-aggfunc-groupby-pandas.ipynb?ref=datascientyst.com) - [pandas.core.groupby.DataFrameGroupBy.describe](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.DataFrameGroupBy.describe.html?ref=datascientyst.com) - [pandas.core.groupby.DataFrameGroupBy.count](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.DataFrameGroupBy.count.html?ref=datascientyst.com) - [pandas.core.groupby.DataFrameGroupBy.size](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.DataFrameGroupBy.size.html?ref=datascientyst.com) - [pandas.core.groupby.DataFrameGroupBy.nunique](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.DataFrameGroupBy.nunique.html?ref=datascientyst.com) - [pandas.core.groupby.SeriesGroupBy.unique](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.SeriesGroupBy.unique.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.first](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.first.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.last](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.last.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.sum](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.sum.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.min](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.min.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.max](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.max.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.mean](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.mean.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.median](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.median.html?ref=datascientyst.com) - [pandas.Series.mode](https://pandas.pydata.org/docs/reference/api/pandas.Series.mode.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.std](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.std.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.var](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.var.html?ref=datascientyst.com) - [pandas.core.groupby.DataFrameGroupBy.mad](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.DataFrameGroupBy.mad.html?ref=datascientyst.com) - [pandas.core.groupby.DataFrameGroupBy.quantile](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.DataFrameGroupBy.quantile.html?ref=datascientyst.com) - [pandas.core.groupby.DataFrameGroupBy.skew](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.DataFrameGroupBy.skew.html?ref=datascientyst.com) - [pandas.core.groupby.GroupBy.sem](https://pandas.pydata.org/docs/reference/api/pandas.core.groupby.GroupBy.sem.html?ref=datascientyst.com) - [pandas.DataFrame.max](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.max.html?ref=datascientyst.com) - [pandas.DataFrame.min](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.min.html?ref=datascientyst.com) - [pandas.DataFrame.median](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.median.html?ref=datascientyst.com) - [pandas.DataFrame.mean](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.mean.html?ref=datascientyst.com) - [pandas.DataFrame.sum](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.sum.html?ref=datascientyst.com) - [pandas.DataFrame.count](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.count.html?ref=datascientyst.com) - [pandas.DataFrame.std](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html?ref=datascientyst.com) - [pandas.DataFrame.mode](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.mode.html?ref=datascientyst.com) - [pandas.DataFrame.var](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.var.html?ref=datascientyst.com) ### How to Group By Multiple Columns in Pandas URL: https://datascientyst.com/use-groupby-multiple-columns-pandas/ Last updated: 2022-10-13T20:05:31.000Z To **group by multiple columns in Pandas DataFrame** can we use the method `groupby()`? We will cover: - **group by multiple columns** - group by several statistical functions - named group by - Series vs DataFrame group by To **group by multiple columns and using several statistical functions** we are going to use next functions: - `groupby()` - `agg()` - `'mean', 'count', 'sum'` ```python df.groupby(['publication', 'date_m']).agg(['mean', 'count', 'sum']) ``` Let's see all the steps in order to find the statistics for each group. ## Step 1: Create sample DataFrame Let's use the following Kaggle Dataset: | title | date | publication | url | | ------------------------------------------------------------------------------------------ | ---------- | -------------------- | ---------------------------------------------------------------------------------------------------------------------------- | | A Round of Applause for Algorithms | 2019-02-18 | Towards Data Science | https://towardsdatascience.com/a-round-of-applause-for-algorithms-3322f6aa1f8e | | The Power of Thank-You Notes: A Simple Way to Make Your Team Happier (And More Productive) | 2019-10-21 | The Startup | https://medium.com/swlh/the-power-of-thank-you-notes-a-simple-way-to-make-your-team-happier-and-more-productive-fc6f2a575de2 | | The Struggle of Modern Day Intrusion Detection Systems | 2019-09-17 | Towards Data Science | https://towardsdatascience.com/the-struggle-of-modern-day-intrusion-detection-systems-50481a6b53c6 | | A Meditation on Stringing Words Together: The National’s “Roman Holiday” | 2019-05-22 | The Startup | https://medium.com/swlh/a-meditation-on-stringing-words-together-the-nationals-roman-holiday-7acbfef7cc02 | | A Sense of Purpose Enables Better Human-Robot Collaboration | 2019-10-14 | Towards Data Science | https://towardsdatascience.com/a-sense-of-purpose-enables-better-human-robot-collaboration-fbe64d0ae913 | You can find the sample data from the repository of the notebook or use the link below to download it. To learn more about reading Kaggle data with Python and Pandas: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) ## Step 2: Group by multiple columns First lets see how to group by a single column in a Pandas DataFrame you can use the next syntax: ```python df.groupby(['publication']) ``` In order to **group by multiple columns** we need to give a list of the columns. **Group by two columns in Pandas**: ```python df.groupby(['publication', 'date_m']) ``` The columns and aggregation functions should be provided as a list to the `groupby` method. ## Step 3: GroupBy SeriesGroupBy vs DataFrameGroupBy The object returned after the groupby of multiple columns depends on the usage of the groups. Let's check it by examples: ```python df.groupby('publication') ``` returns: ``` pandas.core.groupby.generic.DataFrameGroupBy ``` ```python df.groupby(['publication'])['url'] ``` returns: ``` pandas.core.groupby.generic.SeriesGroupBy ``` ```python df.groupby(['publication'])[['url', 'date']] ``` returns: ``` pandas.core.groupby.generic.DataFrameGroupBy ``` If you use a single column after the `groupby` you will get `SeriesGroupBy` otherwise you will have `DataFrameGroupBy`. ## Step 4: Apply multiple statistical functions **Applying multiple aggregation functions to a `groupby` is done by method: `agg`**. Note: that another function `aggregate` exists which and `agg` is an alias for it. The functions can be passed as a list. The **available aggregation functions for group by in Pandas are**: - `count` – non-null values - `min` / \`max – minimum/maximum - `std` – standard deviation - `sum` – sum of values - `mean` / `median` – mean/median - `mode` - `var` Below you can find the syntax in order to **apply multiple statistical functions on multiple pandas columns while grouping by multiple columns**: ```python df.groupby(['publication', 'date_m'])[['claps', 'reading_time']].agg(['mean', 'count', 'sum']) ``` result: | | | claps | reading\_time | | | | | | ------------- | ----------- | ----------- | ------------- | --------- | --------- | ----- | --- | | | | mean | count | sum | mean | count | sum | | publication | date\_m | | | | | | | | Better Humans | 2019-03 | 1710.800000 | 5 | 8554 | 12.000000 | 5 | 60 | | 2019-04 | 3942.250000 | 4 | 15769 | 7.500000 | 4 | 30 | | | 2019-05 | 1187.000000 | 4 | 4748 | 25.000000 | 4 | 100 | | | 2019-06 | 9700.000000 | 1 | 9700 | 5.000000 | 1 | 5 | | | 2019-07 | 1183.333333 | 3 | 3550 | 22.333333 | 3 | 67 | | If you don't provide columns after the grouper than all numeric/datetime columns will be returned: ```python df.groupby(['publication', 'date_m']).agg(['mean']) ``` all columns on which agg functions can be applied are returned: | | | id | claps | reading\_time | date | | ------------- | ----------- | ----------- | ----------- | ------------------- | ------------------- | | | | mean | mean | mean | mean | | publication | date\_m | | | | | | Better Humans | 2019-03 | 5138.200000 | 1710.800000 | 12.000000 | 2019-03-17 04:48:00 | | 2019-04 | 2308.250000 | 3942.250000 | 7.500000 | 2019-04-24 18:00:00 | | | 2019-05 | 3382.000000 | 1187.000000 | 25.000000 | 2019-05-14 06:00:00 | | | 2019-06 | 6326.000000 | 9700.000000 | 5.000000 | 2019-06-27 00:00:00 | | | 2019-07 | 1189.333333 | 1183.333333 | 22.333333 | 2019-07-24 00:00:00 | | Note that the above results have MultiIndex. In order to work with MultiIndex you can check those articles: - [How to Sort MultiIndex in Pandas](https://datascientyst.com/sort-multiindex-pandas/) - [What is a DataFrame MultiIndex in Pandas](https://datascientyst.com/dataframe-multiindex-in-pandas/) ![](https://datascientyst.com/content/images/2021/08/use-groupby-multiple-columns-pandas.png) ## Step 5: Pandas groupby and named aggregations What if you like to **group by multiple columns with several aggregation functions** and would like to have - **named aggregations.** In other words you need to get the next information: - `id` \- mean of this column - `claps` \- `mean`, `count` and `range` You can also name the columns to meet your needs. Finally you can get them without MultiIndex. This is possible by next syntax: ```python df.groupby('publication').agg( id_mean = ('id', 'mean'), claps_mean = ('claps', 'mean'), claps_count = ('claps', 'count'), claps_range = ('claps', lambda x: x.max() - x.min())) ``` this will return: | | id\_mean | claps\_mean | claps\_count | claps\_range | | ----------------------- | ----------- | ----------- | ------------ | ------------ | | publication | | | | | | Better Humans | 3004.285714 | 1827.785714 | 28 | 9660 | | Better Marketing | 2954.578512 | 829.347107 | 242 | 22971 | | Data Driven Investor | 3491.115681 | 95.034704 | 778 | 4100 | | The Startup | 3238.427162 | 303.403815 | 3041 | 38000 | | The Writing Cooperative | 3539.476427 | 372.744417 | 403 | 12600 | ## Conclusion In this article we saw\*\* how to group by multiple columns in Pandas.\*\* We saw how to group by two columns and use different aggregation functions in Pandas DataFrame. ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/groupby/2.use-groupby-multiple-columns-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.agg](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.agg.html?ref=datascientyst.com) - [named-aggregation](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/groupby.html?highlight=filter&ref=datascientyst.com#named-aggregation) ### How to Sort MultiIndex in Pandas URL: https://datascientyst.com/sort-multiindex-pandas/ Last updated: 2021-08-27T20:50:29.000Z This tutorial will show how to **sort MultiIndex in Pandas**. Let's begin by showing the syntax for sorting MultiIndex: ```python .sort_values(by=[('Level 1', 'Level 2')], ascending=False) ``` In order to sort MultiIndex you need to provide all levels which will be used for the sort. Otherwise you will get error like: > ValueError: The column label 'Depth' is not unique. > For a multi-index, the label must be a tuple with elements corresponding to each level. ## Step 1: Create MultiIndex DataFrame Very often multiple aggregation function will end into MultiIndex. It's quite common to sort the MultiIndex which is result of this aggregation. So let's have this DataFrame: | Magnitude Type | Depth | Magnitude | | -------------- | ----- | --------- | | MB | 100.0 | 5.6 | | MWC | 10.0 | 5.5 | | MWW | 21.0 | 6.0 | | MWC | 35.0 | 5.5 | | MWB | 45.0 | 5.6 | For this DataFrame we would like to group by `Magnitude Type` and get the mean, count and sum for columns - `'Depth', 'Magnitude'`. ```python df_multi = df.groupby(['Magnitude Type'])[['Depth', 'Magnitude']].agg(['mean', 'count', 'sum']) ``` This would result into: | | Depth | Magnitude | | | | | | -------------- | --------- | --------- | ---------- | -------- | ----- | -------- | | | mean | count | sum | mean | count | sum | | Magnitude Type | | | | | | | | MB | 81.579365 | 3761 | 306819.990 | 5.682957 | 3761 | 21373.60 | | MD | 21.670000 | 6 | 130.020 | 5.966667 | 6 | 35.80 | | MH | 8.074600 | 5 | 40.373 | 6.540000 | 5 | 32.70 | | ML | 14.158273 | 77 | 1090.187 | 5.814675 | 77 | 447.73 | | MS | 30.142226 | 1702 | 51302.068 | 5.994360 | 1702 | 10202.40 | | MW | 77.034037 | 7722 | 594856.835 | 5.933794 | 7722 | 45820.76 | | MWB | 76.989829 | 2458 | 189241.000 | 5.907282 | 2458 | 14520.10 | | MWC | 66.808213 | 5669 | 378735.760 | 5.858176 | 5669 | 33210.00 | | MWR | 22.445385 | 26 | 583.580 | 5.630769 | 26 | 146.40 | | MWW | 67.568545 | 1983 | 133988.425 | 6.008674 | 1983 | 11915.20 | In the next step we will see how to sort the MultiIndex above. ## Step 2: Find the MultiIndex levels Let's see what is stored as MultiIndex in the DataFrame above. Since we have MultiIndex for the columns we can get the information about the levels by: ```python df_multi.columns ``` result: ``` MultiIndex([( 'Depth', 'mean'), ( 'Depth', 'count'), ( 'Depth', 'sum'), ('Magnitude', 'mean'), ('Magnitude', 'count'), ('Magnitude', 'sum')], ) ``` To get a specific level we can do: ```python df_multi.columns.get_level_values(1) ``` result: ``` Index(['mean', 'count', 'sum', 'mean', 'count', 'sum'], dtype='object') ``` ## Step 3: Sort MultiIndex in Pandas Now let's say that we would like to sort by `mean` which is under `Depth`. From previous step we saw that we need to use: `[('Depth', 'mean')]` for the `by` parameter: ```python df_multi.sort_values(by=[('Depth', 'mean')], ascending=False).head(60) ``` So now the values are sorted by the pair of `Depth - mean`. | | Depth | Magnitude | | | | | | -------------- | --------- | --------- | ---------- | -------- | ----- | -------- | | | mean | count | sum | mean | count | sum | | Magnitude Type | | | | | | | | MB | 81.579365 | 3761 | 306819.990 | 5.682957 | 3761 | 21373.60 | | MW | 77.034037 | 7722 | 594856.835 | 5.933794 | 7722 | 45820.76 | | MWB | 76.989829 | 2458 | 189241.000 | 5.907282 | 2458 | 14520.10 | | MWW | 67.568545 | 1983 | 133988.425 | 6.008674 | 1983 | 11915.20 | | MWC | 66.808213 | 5669 | 378735.760 | 5.858176 | 5669 | 33210.00 | ## Step 4: Sort MultiIndex by multiple levels What if you like to **sort MultiIndex by multiple levels?** In this case you can use the next syntax: ```python df_multi.sort_values(by=[('Depth', 'mean'), ('Depth', 'sum')], ascending=False) ``` This will sort : - first by - ('Depth', 'mean') - then by - ('Depth', 'sum') ## Step 5: Sort MultiIndex by the level number Finally let's say that you prefer to use the number of the level instead of providing a tuple. In this case you can read the level info from Step 2 and use it. For example sorting the MultiIndex by third level will be: `df_multi.columns[2]` \- which is equivalent to `('Depth', 'sum')`: ```python df_multi.sort_values(by=[df_multi.columns[2]], ascending=False).head(5) ``` ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/multiindex/3.sort-multiindex-pandas.ipynb?ref=datascientyst.com) - [Hierarchical indexing (MultiIndex)](https://pandas.pydata.org/docs/user%5Fguide/advanced.html?ref=datascientyst.com#hierarchical-indexing-multiindex) - [MultiIndex / advanced indexing](https://pandas.pydata.org/docs/user%5Fguide/advanced.html?ref=datascientyst.com) - [Sorting a MultiIndex](https://pandas.pydata.org/docs/user%5Fguide/advanced.html?ref=datascientyst.com#sorting-a-multiindex) ### How to replace values with regex in Pandas URL: https://datascientyst.com/replace-values-regex-pandas/ Last updated: 2021-12-02T10:51:50.000Z In this quick tutorial, we'll show **how to replace values with regex in Pandas DataFrame**. There are several options **to replace a value in a column or the whole DataFrame with regex**: 1. Regex replace **string** ```python df['applicants'].str.replace(r'\sapplicants', '') ``` 1. Regex replace **capture group** ```python df['applicants'].replace(to_replace=r"([0-9,\.]+)(.*)", value=r"\1", regex=True) ``` 1. Regex replace **special characters** \- `r'[^0-9a-zA-Z:,\s]+'` \- including spaces ```python df['internship'].str.replace(r'[^0-9a-zA-Z:,]+', '') ``` 1. Regex replace numbers or non-digit characters ```python df['applicants'].str.replace(r'\D+', '') ``` In the next part of the post, you'll see the steps and **practical examples on how to use regex and replace in Pandas**. ## Step 1: Create Sample DataFrame First, let's create a sample 'dirty' data which needs to be cleaned and replaced: ```python import pandas as pd df = pd.read_csv(f'../data/internshala_dataset_raw.csv') df ``` | internship | applicants | stipend | duration | | ---------------------------------------- | --------------------- | ------------------ | -------- | | Mobile App Development | 49 applicants | 15000-25000 /month | 6 Months | | React/React JS | 35 applicants | 5000 /month | 4 Months | | ReactJS Development | Be an early applicant | 5000 /month | 6 Months | | Video Online Course Development (Python) | 63 applicants | 1000 /month | 1 Month | | Internet Of Things (IoT) | 140 applicants | 6000 /month | 6 Months | Data is real and available from Kaggle: [Internship posted in computer science category](https://www.kaggle.com/dhairyatalsania/internshalaindia-raw-dataset?select=internshala%5Fdataset%5Fraw.csv&ref=datascientyst.com). To learn more about reading Kaggle data with Python and Pandas: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) ## Step 2: Replace String Values with Regex in Column Let's start with **replacing string values in column** `applicants`. As you can see the values in the column are mixed. There are two options: ### Replace single string value ```python df['applicants'].str.replace(r'\sapplicants', '', regex=True) ``` The result of this operation will be a Pandas Series: ``` ['49', '35', 'Be an early applicant', '63', '140'] ``` There are some values which we would like to replace too. This will be done in the next section ### Replace multiple string value To **replace multiple values with regex in Pandas** we can use the following syntax: `r'(\sapplicants|Be an early applicant)'` \- where the values to replaced are separated by pipes - `|` ```python df_temp['applicants'].str.replace(r'(\sapplicants|Be an early applicant)', '', regex=True).to_list() ``` result: ``` ['49', '35', '', '63', '140'] ``` ### Replace multiple string values with different replacement Finally for this step let's see how we can **replace multiple values with different values**. Let say that we would like to replace: - `Be an early applicant` with 0 - `49 applicants` with 49 in this case we can use the method `replace` and pass a list of values to be replaced and list of values of the replacement: ```python df_temp['applicants'].replace([r'(\d+) applicants', 'Be an early applicant'],[r'\1',0], regex=True) ``` result: ``` ['49', '35', 0, '63', '140'] ``` In the last example we saw using a capture group which we are going to check in the next step. ## Step 3: Regex replace with capture group In this step we will take a deeper look on **regex and capture groups in Pandas**. They are powerful tool to match a pattern and extract only part of it. Let's say that we would like to match : `63 applicants` but only extract the numbers. In other words, to search for a numeric sequence followed by anything. We are capturing two groups but will keep only the first one - the numbers. ```python df['applicants'].replace(to_replace=r"([0-9,\.]+)(.*)", value=r"\1", regex=True) ``` the result would be: ``` ['49', '35', 'Be an early applicant', '63', '140'] ``` As you can see the code works as expected in case of a match. Otherwise it will keep the value. ## Step 4: Regex replace only special characters What if we would like to clean or remove all **special characters while keeping numbers and letters**. In that case we can use one of the next regex: - `r'[^0-9a-zA-Z:,\s]+'` \- keep numbers, letters, semicolon, comma and space - `r'[^0-9a-zA-Z:,]+'` \- keep numbers, letters, semicolon and comma So the code looks like: ```python df['internship'].str.replace(r'[^0-9a-zA-Z:,\s]+', '', regex=True) ``` Will change: `Internet Of Things (IoT)` to `Internet Of Things IoT` ## Step 5: Regex replace numbers or non-digit characters Now let's check how we can\*\* replace all non digit characters and convert the value to int or remove all numbers from a column\*\*. ### Replace all non numeric symbols and map in case of missing In this example we are going **to replace everything which is not a number with a regex**. In case of a value which doesn't have a number we will map the value to 0. ```python df_temp['applicants'].str.replace(r'\D+', '', regex=True).replace({'':0}).astype('int') ``` This is done on three simple steps: - first replace all non numeric symbols - `str.replace(r'\D+', '', regex=True)` - second - in case of missing numbers - empty string is returned - map the empty string to 0 by `.replace({'':0})` - convert to numeric column ### Replace all numbers from Pandas column To **replace all numbers from a given column** you can use the next syntax: ```python df['applicants'].replace(to_replace=r"\d+", value=r" ", regex=True) ``` result: ``` [' applicants', ' applicants', 'Be an early applicant', ' applicants', ' applicants'] ``` ## Step 6: Regex replace all values in DataFrame Finally let's find out how to **replace values in the whole DataFrame - all columns and rows** \- with a single line of code: To replace all numbers from a DataFrame use method `replace` \- directly on the DataFrame: ```python df_temp.replace(to_replace=r"\d+", value=r" ", regex=True) ``` result: | internship | applicants | stipend | duration | | ---------------------------------------- | --------------------- | --------- | -------- | | Mobile App Development | applicants | \- /month | Months | | React/React JS | applicants | /month | Months | | ReactJS Development | Be an early applicant | /month | Months | | Video Online Course Development (Python) | applicants | /month | Month | | Internet Of Things (IoT) | applicants | /month | Months | ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/regex/1.replace-values-regex-pandas.ipynb?ref=datascientyst.com) - [pandas.DataFrame.replace](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.replace.html?ref=datascientyst.com) - [pandas.Series.str.replace](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.replace.html?ref=datascientyst.com) - [pandas.Series.str.extract](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.extract.html?ref=datascientyst.com) - [How to Add New Column Based on List of Keywords in Pandas DataFrame](https://datascientyst.com/add-new-column-list-keywords-pandas-dataframe/) ### How to Compare Titles and URLs in Pandas URL: https://datascientyst.com/compare-titles-urls-pandas/ Last updated: 2021-12-02T10:52:07.000Z In this tutorial, I'll show you how to compare titles and URLs with Pandas and Python. We are going to use several Pandas functions and techniques in order to find problematic data. To start, here is the initial Data and what is our final goal: | title | url | | --------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | | A Beginner’s Guide to Word Embedding with Gensim Word2Vec Model | https://towardsdatascience.com/a-beginners-guide-to-word-embedding-with-gensim-word2vec-model-5970fa56cc92 | | Hands-on Graph Neural Networks with PyTorch & PyTorch Geometric | https://towardsdatascience.com/hands-on-graph-neural-networks-with-pytorch-pytorch-geometric-359487e221a8 | | How to Use ggplot2 in Python | https://towardsdatascience.com/how-to-use-ggplot2-in-python-74ab8adec129 | | Databricks: How to Save Files in CSV on Your Local Computer | https://towardsdatascience.com/databricks-how-to-save-files-in-csv-on-your-local-computer-3d0c70e6a9ab | | A Step-by-Step Implementation of Gradient Descent and Backpropagation | https://towardsdatascience.com/a-step-by-step-implementation-of-gradient-descent-and-backpropagation-d58bda486110 | Usually the slug of the URL is composed from the title after several transformations. Below you can find example pair of URL and title: - [https://towardsdatascience.com/a-beginners-guide-to-word-embedding-with-gensim-word2vec-model-5970fa56cc92](https://towardsdatascience.com/a-beginners-guide-to-word-embedding-with-gensim-word2vec-model-5970fa56cc92?ref=datascientyst.com) - A Beginner’s Guide to Word Embedding with Gensim Word2Vec Model Our goal is to find all articles which have mismatch between the title and the slug like: - Faster Training for Efficient CNNs - faster-training-for-efficient-cnns - [https://towardsdatascience.com/faster-training-of-efficient-cnns-657953aa080](https://towardsdatascience.com/faster-training-of-efficient-cnns-657953aa080?ref=datascientyst.com) The difference above is `for` vs `of`. In the next section, I’ll review the steps to compare and validate the title against the URL of a given article. ## Step 1: Read the data and install slugify In this tutorial we are going to use data available from Kaggle. You can find it also in the repository related to the notebook(in Resources): ```python import pandas as pd df = pd.read_csv('../data/medium_data.csv.zip') df ``` If you like to learn more about how to read Kaggle as a Pandas DataFrame check this article: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) We need to install Python library called: `slugify`: ```python pip install slugify ``` ## Step 2: Convert title to slug with slugify Next step is to prepare the titles in order to compare them with the URLs. For this purpose we are going to use library `slugify`: ```python from slugify import slugify df['title_to_url'] = df['title'].fillna('').apply(lambda x: slugify(x)) ``` After this operation we will have the next data: | title | url | title\_to\_url | | --------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- | | A Beginner’s Guide to Word Embedding with Gensim Word2Vec Model | https://towardsdatascience.com/a-beginners-guide-to-word-embedding-with-gensim-word2vec-model-5970fa56cc92 | a-beginners-guide-to-word-embedding-with-gensim-word2vec-model | | Hands-on Graph Neural Networks with PyTorch & PyTorch Geometric | https://towardsdatascience.com/hands-on-graph-neural-networks-with-pytorch-pytorch-geometric-359487e221a8 | hands-on-graph-neural-networks-with-pytorch-pytorch-geometric | | How to Use ggplot2 in Python | https://towardsdatascience.com/how-to-use-ggplot2-in-python-74ab8adec129 | how-to-use-ggplot2-in-python | | Databricks: How to Save Files in CSV on Your Local Computer | https://towardsdatascience.com/databricks-how-to-save-files-in-csv-on-your-local-computer-3d0c70e6a9ab | databricks-how-to-save-files-in-csv-on-your-local-computer | | A Step-by-Step Implementation of Gradient Descent and Backpropagation | https://towardsdatascience.com/a-step-by-step-implementation-of-gradient-descent-and-backpropagation-d58bda486110 | a-step-by-step-implementation-of-gradient-descent-and-backpropagation | ## Step 3: Find all rows where the slugified titles is not in the URL In this step we are going to compare the slugified titles and the URLs. This is a pretty generic comparison and we might miss some false positives. We are going to create a new column - `bad_title` where all values will be set to `False`. Only if the slugified title is not part of the URL then we are going to set it to `True`: ```python df['bad_title'] = False for row in df[df.title.notna()].iterrows(): title_to_url_temp = row[1].title_to_url.lower() title_temp = row[1].title.replace('-', '').replace(' ', '').lower() if not title_to_url_temp in row[1]['url']: df.loc[row[0], 'bad_title'] = True ``` In total - 1051 results like: | title | url | title\_to\_url | | ----------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | | What I Learned from (Two-time) Kaggle Grandmaster Abhishek Thakur | https://towardsdatascience.com/what-i-learned-from-abhishek-thakur-4b905ac0fd55 | em-class-markup-em-markup-h3-em-what-i-learned-from-two-time-kaggle-grandmaster-abhishek-thakur-em | | Faster Training for Efficient CNNs | https://towardsdatascience.com/faster-training-of-efficient-cnns-657953aa080 | faster-training-for-efficient-cnns | | Buyers beware, fake product reviews are plaguing the internet. How Machine Learning can help to spot them. | https://towardsdatascience.com/buyers-beware-fake-product-reviews-are-plaguing-the-internet-cfc599c42b6b | buyers-beware-fake-product-reviews-are-plaguing-the-internet-how-machine-learning-can-help-to-spot-them | You can use different techniques in order to get better results like: - replace hyphens from both sides - convert to lower case ## Step 4: Extract the slug from the URL and compare to title In this step we are going to extract the slug from the URL. This is not a simple operation and might depend on many factors. We can't replace the domain or extract the slug with simple split: ```python df['url'].str.split('/', expand = True) ``` Result: | 0 | 1 | 2 | 3 | 4 | | ------ | - | ---------------------- | ------------------------------------------------------------------------- | ---------------------------------------------------------------------- | | https: | | medium.com | swlh | how-to-go-from-zero-to-hero-through-personal-branding-e5edab472dfa | | https: | | medium.com | swlh | five-augmented-reality-uses-that-solve-real-life-problems-9e7ff77be858 | | https: | | towardsdatascience.com | everything-you-need-to-know-about-autoencoders-in-tensorflow-b6a63e8255f0 | None | | https: | | medium.com | datadriveninvestor | electric-company-cars-get-a-tax-incentive-in-germany-a0c207da0bb9 | | https: | | uxdesign.cc | ux-is-not-a-role-it-is-a-title-b072aff1c98c | None | As you can see there are multiple domains and formats. What we are going to do is extracting the slug by chaining several Pandas operations: ```python df['url'].str.rsplit('/').str[-1].str.rsplit('-', 1, expand=True)[0] ``` \*\*How does it work? \*\* Starting with - [https://towardsdatascience.com/how-to-use-ggplot2-in-python-74ab8adec129](https://towardsdatascience.com/how-to-use-ggplot2-in-python-74ab8adec129?ref=datascientyst.com) We are using the method `rsplit` in order to split on `/`. ``` [https:, , towardsdatascience.com, how-to-use-ggplot2-in-python-74ab8adec129] ``` Then adding `str[-1]` in order to get the part after the last backslash: ``` how-to-use-ggplot2-in-python-74ab8adec129 ``` In order to get rid of that final identifier - `-74ab8adec129` we will use `rsplit` again - `.rsplit('-', 1, expand=True)[0]` Now we can compare the slug from the URL with the modified title: ```python df[df['url'].str.rsplit('/').str[-1].str.rsplit('-', 1, expand=True)[0] != df['title_to_url']] ``` The result is pretty similar to the one from Step 3 - but we have more results - 1119. But there is a difference between both methods. This step is more accurate and will report cases like: - Title: `Living as an Empath — w` - URL: `https://medium.com/swlh/living-as-an-empath-when-you-feel-everything-and-nobody-seems-to-understand-65fecce2aa03` - title\_to\_url: `living-as-an-empath-w` while Step 3 will miss this result since the `title_to_url` is part of the URL. ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/tutorials/1.compare-titles-urls-pandas.ipynb?ref=datascientyst.com) - [pypi - slugify](https://pypi.org/project/slugify/?ref=datascientyst.com) ### Progress Bar for Merge Or Concat Operation With tqdm in Pandas URL: https://datascientyst.com/progress-bar-merge-concat-operation-tqdm-pandas/ Last updated: 2021-08-26T10:55:47.000Z Need a **progress bar for Pandas concat, merge or join operations**? If so, you may use the workaround described in this article. At the moment of writing there's no simple solution for Pandas and `tqdm` in order to track progress for `merge` or `concat`. In order to get the progress during those operations we will use Dask. ## Step 1: Install Dask and TQDM `Dask`\`tqdm\` libraries can be installed by: ```python pip install tqdm pip install dask ``` and upgraded by: ```python pip install tqdm -U pip install dask -U ``` ## Step 2: Create and convert Pandas DataFrames to Dask First we are going to **create two medium sized DataFrames in Pandas** with random numbers from 0 to 700. Then we are going to **convert them to Dask DataFrames.** ```python import pandas as pd import numpy as np from tqdm import tqdm import dask.dataframe as dd n = 450000 maxa = 700 df1 = pd.DataFrame({'lkey': np.random.randint(0, maxa, n),'lvalue': np.random.randint(0,int(1e8),n)}) df2 = pd.DataFrame({'rkey': np.random.randint(0, maxa, n),'rvalue': np.random.randint(0, int(1e8),n)}) sd1 = dd.from_pandas(df1, npartitions=3) sd2 = dd.from_pandas(df2, npartitions=3) ``` ## Step 3: Add progress bar for merge on two DataFrames Finally we are going to use the **Dask progress bar in order to track the progress on merging of two DataFrames**. The are two options available - use context `with TqdmCallback(desc="compute")` ```python from tqdm.dask import TqdmCallback from dask.diagnostics import ProgressBar ProgressBar().register() with TqdmCallback(desc="compute"): sd1.merge(sd2, left_on='lkey', right_on='rkey').compute() ``` or use it globally: ```python # or use callback globally cb = TqdmCallback(desc="global") cb.register() sd1.merge(sd2, left_on='lkey', right_on='rkey').compute() ``` result: ``` [ ] | 0% Completed | 0.0s global: 0%| | 0/31 [00:00 OutOfBoundsDatetime: Out of bounds nanosecond timestamp If so, I'll show you the reason for the problem, how to investigate it and how to solve it. Normally, If you want to convert a string to date with `pd.to_datetime`, you have many options and important details. Some of them can be found in the articles below: - [How to Fix Pandas to\_datetime: Wrong Date and Errors By John D K in How To Guides • 2 months ago](https://datascientyst.com/how-to-fix-pandas-to%5Fdatetime-wrong-date-and-errors/) - [How to convert month number to month name in Pandas DataFrame](https://datascientyst.com/convert-month-number-to-month-name-pandas-dataframe/) ## Step 1: What is "OutOfBoundsDatetime: Out of bounds nanosecond timestamp" To start let's explain what is error: > OutOfBoundsDatetime: Out of bounds nanosecond timestamp Pandas uses NumPy 'datetime64' and 'timedelta64' dtypes in order to add features around Time Series. This is related to limitations as follows - the range of dates in limited in next interval: ```python pd.Timestamp.min ``` result: ``` Timestamp('1677-09-21 00:12:43.145225') ``` ```python pd.Timestamp.max ``` result: ``` Timestamp('2262-04-11 23:47:16.854775807') ``` Anything which is **outside this date range will raise the error: `OutOfBoundsDatetime: Out of bounds nanosecond timestamp`** So all the code examples below will raise the error: ```python pd.to_datetime('Jun 1, 1111') pd.to_datetime('7 1') pd.to_datetime('Jun 1') ``` ## Step 2: Analyse error "OutOfBoundsDatetime: Out of bounds nanosecond timestamp" In this step you can learn how to analyse the error and find where the problem is. Suppose we have a DataFrame like the one below: ```python import pandas as pd data = { "company":{"0":"Bad Pandas","2":"Clever Fox","4":"Max Wolf","6":"Rage Raycon","8":"Massive Shark"}, "date":{"0":"Jul 16, 2020","2":"Jul 2, 2021","4":"Jun 27, 2019","6":"Jun 13","8":"May 17"}, "sales":{"0":"17","2":"27","4":"202","6":"33","8":"29"}} df = pd.DataFrame(data) ``` | company | date | sales | | ------------- | ------------ | ----- | | Bad Pandas | Jul 16, 2020 | 17 | | Clever Fox | Jul 2, 2021 | 27 | | Max Wolf | Jun 27, 2019 | 202 | | Rage Raycon | Jun 13 | 33 | | Massive Shark | May 17 | 29 | If we try to convert column `date` to a datetime we will end with error: ```python pd.to_datetime(df['date']) ``` output: ``` Out of bounds nanosecond timestamp: 1-06-13 00:00:00 ``` In this case it might be obvious where the problem is: `Jun 13` but in some cases you will have thousands out of millions which will need a fix. To find the records which are causing issues follow: **1) Create a new column for date 'date2' with:** ```python df['date2'] = pd.to_datetime(df['date'], errors = 'coerce') ``` **2) Find all problematic dates** ```python df[df['date2'].isna()] ``` result: ``` Jun 13 May 17 ``` **3) Correct the problematic records** In my case the problem is that dates for the current year are missing the year. So in order to fix that problem we can use: ```python df_temp = df[df['date2'].isna()] df.loc[df_temp.index, 'date']= df_temp['date'] + ', 2021' ``` **4) Convert all dates to datetime with: pd.to\_datetime** The final step is to convert the dates as intended initially: ```python pd.to_datetime(df['date']) ``` ## Step 3: Fix and workarounds for "OutOfBoundsDatetime: Out of bounds nanosecond timestamp" There are several options in order to workaround the error. You can find two of them explained below: 1. Parameter - errors = 'ignore' will convert all dates which are OK and the rest will remain unchanged ```python pd.to_datetime(df['date'], errors = 'ignore') ``` result: ``` 0 Jul 16, 2020 2 Jul 2, 2021 4 Jun 27, 2019 6 Jun 13 8 May 17 ``` 1. Parameter - errors = 'coerce' will convert all dates which are OK and the rest will be NaT ```python pd.to_datetime(df['date'], errors = 'coerce') ``` result: ``` 0 2020-07-16 2 2021-07-02 4 2019-06-27 6 NaT 8 NaT ``` ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/datetime/4.outofboundsdatetime-out-of-bounds-nanosecond-timestamp-pandas-pd-to%5Fdatetime.ipynb?ref=datascientyst.com) - [1pandas.to\_datetime](https://pandas.pydata.org/docs/reference/api/pandas.to%5Fdatetime.html?ref=datascientyst.com) - [Pandas Timestamp Limitations](http://pandas-docs.github.io/pandas-docs-travis/user%5Fguide/timeseries.html?ref=datascientyst.com#timestamp-limitations) - [Time Series / Date functionality](http://pandas-docs.github.io/pandas-docs-travis/user%5Fguide/timeseries.html?ref=datascientyst.com) ### How to apply function to multiple columns in Pandas URL: https://datascientyst.com/apply-function-multiple-columns-pandas/ Last updated: 2021-09-11T21:17:46.000Z You can use the following code to **apply a function to multiple columns in a Pandas DataFrame**: ```python def get_date_time(row, date, time): return row[date] + ' ' +row[time] df.apply(get_date_time, axis=1, date='Date', time='Time') ``` For applying function to single column and performance optimization on `apply` check - [How to apply function to single column in Pandas](https://datascientyst.com/apply-function-to-single-column-in-pandas/) Next, you'll see several examples on how to **apply a function to two and more columns in Pandas.** The DataFrame below is available from Kaggle: | Date | Latitude | Longitude | Depth | Type | | ---------- | --------- | ---------- | ----- | ---------- | | 12/24/2016 | \-5.1460 | 153.5166 | 30.00 | Earthquake | | 12/25/2016 | \-43.4029 | \-73.9395 | 38.00 | Earthquake | | 12/25/2016 | \-43.4810 | \-74.4771 | 14.93 | Earthquake | | 12/27/2016 | 45.7192 | 26.5230 | 97.00 | Earthquake | | 12/28/2016 | 38.3754 | \-118.8977 | 10.80 | Earthquake | You can download it from Kaggle or read it with Python - [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) ## Option 1: Apply function to two columns in Pandas DataFrame Suppose you would like to create a new column with the city based on the pair: **Latitude** and **Longitude**. For this purpose we will define new function `geo_rev(x)` which will be applied on columns and will return the city for each row: ```python import geocoder def geo_rev(x): g = geocoder.osm([x['Latitude'], x['Longitude']], method='reverse').json if g: return g.get('country') else: return 'no country' df.apply(geo_rev, axis=1) ``` Function `apply` takes argument `axis=1` which can be described as: - 0 or 'index': apply function to each column. - 1 or 'columns': apply function to each row. The function receives all values from the current row and they can be accessed by: `x['Latitude']` To create a new column after applying a function we can use: ```python df['country'] = df.apply(geo_rev, axis=1) ``` ## Option 2: Apply function to multiple columns with parameters If you need to **apply a function to DataFrame and pass parameters to the function** at the same time then you can use the following syntax: ```python def get_date_time(row, date, time): return row[date] + ' ' +row[time] df.apply(get_date_time, axis=1, date='Date', time='Time') ``` There's no limit on the number of parameters. ## Option 3: Apply function with lambda and multiple columns In this example we are going to use **method `apply` and `lambda` in order to apply function to several columns.** Again we are going to convert `Latitude` and `Longitude` to country by applying function: ```python import pandas as pd def geo_rev(lat, lon): g = geocoder.osm([lat, lon], method='reverse').json if g: return g.get('country') else: return 'no country' df.apply(lambda x: geo_rev(x['Latitude'], x['Longitude']), axis=1) ``` result is: ``` 23402 Papua Niugini 23403 Chile 23404 Chile 23405 România 23406 United States 23407 United States 23408 United States 23409 日本 23410 Indonesia 23411 no country ``` ## Option 4: Select and apply function to multiple columns You can **select several columns from a Pandas DataFrame and apply function to them by**: ```python def geo_rev(lat, lon, mag): g = geocoder.osm([lat, lon], method='reverse').json if g: return g.get('country') + ' ' + str(mag) else: return 'no country ' df[['Latitude', 'Longitude', 'Magnitude']].apply(lambda x: geo_rev(*x), axis=1) ``` result of this operation is: ``` 23402 Papua Niugini 5.8 23403 Chile 7.6 23404 Chile 5.6 ``` ## Option 5: Apply function to multiple columns without using apply Finally let's see an alternative solution to apply a function to several columns but without the method `apply`. This can be achieved by using a combination of `list` and `map`. This technique is **much faster than using Pandas `apply`**: ```python def geo_rev(lat, lon): g = geocoder.osm([lat, lon], method='reverse').json if g: return g.get('country') else: return 'no country ' list(map(geo_rev, df['Latitude'], df['Longitude'])) ``` The advantage of this approach is the speed as we can see in the comparison below for a small dataset: - `%timeit list(map(` \- 12 µs per loop - `%timeit df.apply(` \- 760 µs per loop ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/apply/1.apply-function-multiple-columns-pandas.ipynb?ref=datascientyst.com) - [Reverse Geocoding - Latitude/ Longitude to City/Country - Python and Pandas](https://datascientyst.com/reverse-geocoding-latitude-longitude-city-country-python-pandas/) - [1pandas.DataFrame.apply](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.apply.html?ref=datascientyst.com) - [pandas.Series.apply](https://pandas.pydata.org/docs/reference/api/pandas.Series.apply.html?ref=datascientyst.com) - [pandas.DataFrame.applymap](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.applymap.html?ref=datascientyst.com) ### How to Filter DataFrame by Date in Pandas URL: https://datascientyst.com/filter-by-date-pandas-dataframe/ Last updated: 2021-12-02T10:54:02.000Z Here are several approaches to **filter rows in Pandas DataFrame by date:** **1) Filter rows between two dates** ```python df[(df['date'] > '2019-12-01') & (df['date'] < '2019-12-31')] ``` **2) Filter rows by date in index** ```python df2.loc['2019-12-01':'2019-12-31'] ``` **3) Filter rows by date with Pandas query** ```python df.query('20191201 < date < 20191231') ``` In the next section, you'll see several examples of how to apply the above approaches using simple examples. As a bonus you can find information **how to filter rows per month, week, year, quarter etc.** The DataFrame for the examples below is available from Kaggle. | date | title | | ---------- | --------------------------------------------------------------------- | | 2019-05-30 | A Beginner’s Guide to Word Embedding with Gensim Word2Vec Model | | 2019-05-30 | Hands-on Graph Neural Networks with PyTorch & PyTorch Geometric | | 2019-05-30 | How to Use ggplot2 in Python | | 2019-05-30 | Databricks: How to Save Files in CSV on Your Local Computer | | 2019-05-30 | A Step-by-Step Implementation of Gradient Descent and Backpropagation | If you like to learn more about how to read Kaggle as a Pandas DataFrame check this article: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) Note: For all examples we need to convert column `date` to datetime with method - `to_datetime`: ```python df['date'] = pd.to_datetime(df['date']) ``` ## Option 1: Filter DataFrame by date in Pandas To start with a simple example, let's filter the DataFrame by two dates: - '2019-12-01' - '2019-12-31' We would like to **get all rows which have date between those two dates.** So filtering the rows which meet the above requirement can be done: ```python df[(df['date'] > '2019-12-01') & (df['date'] < '2019-12-31')] ``` result is: ``` 614 rows ``` Note: The format is YYYY-MM-DD Note 2: the filtering above is exclusive. If you like to get inclusive filtering just add `=`: ```python df[(df['date'] >= '2019-12-01') & (df['date'] <= '2019-12-31')] ``` ## Option 2: Filter DataFrame by date using the index For this example we will change the original index of the DataFrame in order to have a index which is a date: ```python df = df.set_index('date') ``` Now having a **DataFrame with index which is a datetime we can filter the rows by**: ```python df.loc['2019-12-01':'2019-12-31'] ``` the result is the same: ``` 614 rows ``` ### 2.1 MultiIndex filtering by date The example below shows how to do **filtering when you have a date in hierarchical index**: ```python import datetime df3 = df.set_index(['publication', 'date']) df3[(df3.index.get_level_values('date') > '2019-12-01')] ``` result: ``` 614 rows ``` Find all article written on Sunday-s: ```python df3[df3.index.get_level_values('date').dayofweek == 6] ``` ## Option 3: Filter rows by date and pd.Timestamp You can use `pd.Timestamp` in order to construct your dates and compare the value with each row. The syntax for creating a date with Pandas is: ```python pd.Timestamp(2019,12,1) ``` So the comparison will be: ```python df[(df['date'] > pd.Timestamp(2019,12,1)) & (df['date'] < pd.Timestamp(2019,12,31))] ``` You can pass a string like in option 1 and Pandas will do the conversion for you. This can be useful when conversion is not explicit or you like to have control on the format. ## Option 4: Pandas filter rows by date with Query Pandas offers a simple way to query rows by method `query`. The syntax is minimalistic and self-explanatory: ```python df.query('20191201 < date < 20191231') ``` result: ``` 614 rows ``` ## Option 5: Pandas filter rows by day, month, week, quarter etc Finally let's check several useful and frequently used filters. ### Filter rows by day ```python df[df['date'].dt.strftime('%Y-%m-%d') == '2019-12-30'] ``` result: ``` 165 rows ``` ### Filter rows by day of the week ```python df[df['date'].dt.dayofweek == 6] ``` result: ``` 609 rows ``` ### Filter rows by month and year ```python df[df['date'].dt.to_period('M') == '2019-12'] ``` result: ``` 614 rows ``` ### Filter rows by month ```python df[df['date'].dt.month == 12] ``` result: ``` 614 rows ``` ### Filter rows by quarter ```python df[df['date'].dt.to_period('Q') == '2019Q4'] ``` or ```python df[df['date'].dt.to_period('Q') == '2019-12'] ``` result: ``` 1915 rows ``` ### Filter by year ```python df[df['date'].dt.to_period('Y') == '2019'] ``` result: ``` 6508 rows ``` ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/filter/1.filter-by-date-pandas-dataframe.ipynb?ref=datascientyst.com) - [Time series / date functionality](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/timeseries.html?ref=datascientyst.com) - [pandas.DataFrame.query](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.query.html?ref=datascientyst.com) - [Indexing and selecting data](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/indexing.html?ref=datascientyst.com) ### Reverse Geocoding - Latitude/ Longitude to City/Country - Python and Pandas URL: https://datascientyst.com/reverse-geocoding-latitude-longitude-city-country-python-pandas/ Last updated: 2021-08-19T20:11:52.000Z Do you need to do **reverse geocoding in Python and Pandas**. Do you need to **convert pairs of latitude and longitude to city, country, address coordinates**? If so in this article you can find two different approaches: - `geocoder` ```python geocoder.osm([51.5074, 0.1278], method='reverse') ``` - `geopy` ```python geolocator.reverse("48.8588443, 2.2943506") ``` In the next section you can find more example usage of both methods. If you need to plot a map from geo coordinates like please check: [How to plot latitude and longitude from Pandas DataFrame in Python](https://datascientyst.com/plot-latitude-longitude-pandas-dataframe-python/) ## Step 1: Install required libraries - geopy and geocoder First you need to install two libraries `geocoder` and `geopy`: ```bash pip install geopy pip install geocoder ``` Both libraries offer providers such as Google & Bing which have geocoding services. `geopy` and `geocoder` help to locate the coordinates of addresses, cities, countries, and landmarks across the globe using third-party geocoders and other data sources. ## Step 2: Convert latitude and longitude to city/country with geopy First we will do reverse geocoding with `geopy`. We are going to retrieve city and state/country about: `51.5074, 0.1278`. We are going to use `user_agent="http"` otherwise errors will be raised. User\_Agent is an http request header that is sent with each request. The code: ```python from geopy.geocoders import Nominatim geolocator = Nominatim(user_agent="http") location = geolocator.reverse("51.5074, 0.1278") print(location.address) ``` this returns: ``` Fleming Way, London Borough of Bexley, London, Greater London, England, SE28 8NS, United Kingdom ``` To return the city and the country you can use the next code format: ```python location.raw.get('address').get('city') ``` output: ``` 'London' ``` You can find all fields returned by: ```python location.raw ``` which are: ```json {'place_id': 85508570, 'licence': 'Data © OpenStreetMap contributors, ODbL 1.0. https://osm.org/copyright', 'osm_type': 'way', 'osm_id': 5181183, 'lat': '51.5075415', 'lon': '0.1280774', 'display_name': 'Fleming Way, London Borough of Bexley, London, Greater London, England, SE28 8NS, United Kingdom', 'address': {'road': 'Fleming Way', 'suburb': 'London Borough of Bexley', 'city': 'London', 'state_district': 'Greater London', 'state': 'England', 'postcode': 'SE28 8NS', 'country': 'United Kingdom', 'country_code': 'gb'}, 'boundingbox': ['51.5071355', '51.5079528', '0.1280173', '0.1292619']} ``` ## Step 3: Reverse Geocoding with geocoder In this step we are going to use `geocoder` in order to return the information about the: - country - city - zip code / postcode - address - start First we will start we getting the city: ```python import geocoder g = geocoder.osm([51.5074, 0.1278], method='reverse') g.json['city'] ``` this will return: ``` 'London' ``` You can find the whole json result below: ```json {'accuracy': 0.001, 'address': 'Fleming Way, London Borough of Bexley, London, Greater London, England, SE28 8NS, United Kingdom', 'bbox': {'northeast': [51.5079528, 0.1292619], 'southwest': [51.5071355, 0.1280173]}, 'city': 'London', 'confidence': 10, 'country': 'United Kingdom', 'country_code': 'gb', 'importance': 0.001, 'lat': 51.5073599, 'lng': 0.1283349, 'ok': True, 'osm_id': 5181183, 'osm_type': 'way', 'place_id': 85508570, 'place_rank': 26, 'postal': 'SE28 8NS', 'quality': 'residential', 'raw': {'place_id': 85508570, 'licence': 'Data © OpenStreetMap contributors, ODbL 1.0. https://osm.org/copyright', 'osm_type': 'way', 'osm_id': 5181183, 'boundingbox': ['51.5071355', '51.5079528', '0.1280173', '0.1292619'], 'lat': '51.5073599', 'lon': '0.1283349', 'display_name': 'Fleming Way, London Borough of Bexley, London, Greater London, England, SE28 8NS, United Kingdom', 'place_rank': 26, 'category': 'highway', 'type': 'residential', 'importance': 0.001, 'address': {'road': 'Fleming Way', 'suburb': 'London Borough of Bexley', 'city': 'London', 'state_district': 'Greater London', 'state': 'England', 'postcode': 'SE28 8NS', 'country': 'United Kingdom', 'country_code': 'gb'}}, 'region': 'England', 'state': 'England', 'status': 'OK', 'street': 'Fleming Way', 'suburb': 'London Borough of Bexley', 'type': 'residential'} ``` ## Step 4: Batch Reverse Geocoding with Pandas Finally let's cover how to reverse geocode with Pandas. We are going to use apply and pass a function which will get as input `Latitude` and `Longitude` and will return new column with information for the city: ```python import geocoder def geo_rev(x): g = geocoder.osm([x.Latitude, x.Longitude], method='reverse').json if g: return g.get('country') else: return 'no country' df[['Latitude', 'Longitude']].apply(geo_rev, axis=1) ``` result: ``` United States United States 日本 Indonesia no country ``` In case of an error or point which is not on the land we will return **no country**. ## Step 5: Errors on geocoding with Python If you try to get result for bad coordinates you will get errors like: for `geocoder` > ValueError: Coords are not within the world's geographical boundary or for `geopy` > ValueError: Must be a coordinate pair or Point ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/geocoding/2.reverse-geocoding-latitude-longitude-city-country-python-pandas.ipynb?ref=datascientyst.com) - [geocoder](https://pypi.org/project/geocoder/?ref=datascientyst.com) - [geopy](https://pypi.org/project/geopy/?ref=datascientyst.com) - [Python Geocoder](https://github.com/DenisCarriere/geocoder?ref=datascientyst.com) ### Plot Latitude and Longitude from Pandas DataFrame in Python URL: https://datascientyst.com/plot-latitude-longitude-pandas-dataframe-python/ Last updated: 2022-10-29T21:12:52.000Z Need to **plot latitude and longitude from Pandas DataFrame in Python**? If so, you may use the following libraries to do so: - `geopandas` - `shapely` - `matplotlib` \- optional - if the map is not displayed - `plotly` \- alternative solution Below you can find working example and all the steps in order to **convert pairs of latitude and longitude to a world map**. ## Step 1: Install required libraries - geopandas In order to use the code below you need the latest Python. You need to install several libraries with: ```bash pip install geopandas pip install Shapely pip install matplotlib ``` The `matplotlib` library is needed for Jupyter Notebook if the plot is not shown. It's a good idea to update all libraries to the latest versions. ## Step 2: Create DataFrame with geocoding data For this example we are going to use data from [Significant Earthquakes, 1965-2016](https://www.kaggle.com/usgs/earthquake-database?select=database.csv&ref=datascientyst.com). You can download it from Kaggle or read it with Python - [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) You can find the data below: ```python import pandas as pd df = pd.read_csv(f'../data/earthquakes_1965_2016_database.csv.zip') df ``` data: | Date | Latitude | Longitude | Depth | Type | | ---------- | -------- | --------- | ----- | ---------- | | 01/02/1965 | 19.246 | 145.616 | 131.6 | Earthquake | | 01/04/1965 | 1.863 | 127.352 | 80.0 | Earthquake | | 01/05/1965 | \-20.579 | \-173.972 | 20.0 | Earthquake | | 01/08/1965 | \-59.076 | \-23.557 | 15.0 | Earthquake | | 01/09/1965 | 11.938 | 126.427 | 15.0 | Earthquake | ## Step 3: Plot latitude and longitude to world map Finally we are going to plot the pairs latitude and longitude on the world map. The points will be paint in red color while to earth will be blue: ```python world = gpd.read_file(gpd.datasets.get_path('naturalearth_lowres')) gdf.plot(ax=world.plot(figsize=(15, 15)), marker='o', color='red', markersize=15); ``` The result is: ![plot-latitude-longitude-pandas-dataframe-python](https://datascientyst.com/content/images/2022/10/plot-latitude-longitude-pandas-dataframe-python.webp) ## Step 4: Plot latitude and longitude to interactive map plus hover with plotly As an alternative solution you can use library **`plotly` to draw a map from latitude and longitude**. You can find the code for it below: ```python import plotly.express as px import pandas as pd fig = px.scatter_geo(df,lat='Latitude',lon='Longitude', hover_name="Magnitude") fig.update_layout(title = 'Significant Earthquakes, 1965-2016', title_x=0.5) fig.show() ``` The advantage of these method is that **plots an interactive map which alows you to hover on the points**. You can also display important information plus coordinates like: - name - latitude - longitude or zoom in/out: ![plot-latitude-longitude-pandas-dataframe-python-interactive-map](https://datascientyst.com/content/images/2022/10/plot-latitude-longitude-pandas-dataframe-python-interactive-map.webp) ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/geocoding/1.plot-latitude-longitude-pandas-dataframe-python.ipynb?ref=datascientyst.com) - [geopandas](https://pypi.org/project/geopandas/?ref=datascientyst.com) - [Shapely](https://pypi.org/project/Shapely/?ref=datascientyst.com) - [matplotlib](https://pypi.org/project/matplotlib/?ref=datascientyst.com) ### How to convert month number to month name in Pandas DataFrame URL: https://datascientyst.com/convert-month-number-to-month-name-pandas-dataframe/ Last updated: 2021-12-02T10:54:29.000Z You can use the following code to **convert the month number to month name in Pandas.** Full name: ```python df['date'].dt.month_name() ``` 3 letter abbreviation of the month: ```python df['date'].dt.month_name().str[:3] ``` Next, you'll see example and steps to get the month name from number: ## Step 1: Read a DataFrame and convert string to a DateTime For example, let's have a DataFrame with a date column inside which is stored a string. DataFrame and steps are explained here - [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) or if you use the notebook: ```python import pandas as pd df = pd.read_csv('../data/medium_data.csv.zip') df ``` In case of a date which is stored as a string you can convert it to a datetime by: ```python df['date'] = pd.to_datetime(df['date']) ``` the data is shown below: | title | date | claps | | --------------------------------------------------------------------- | ---------- | ----- | | A Beginner’s Guide to Word Embedding with Gensim Word2Vec Model | 2019-05-30 | 850 | | Hands-on Graph Neural Networks with PyTorch & PyTorch Geometric | 2019-05-30 | 1100 | | How to Use ggplot2 in Python | 2019-05-30 | 767 | | Databricks: How to Save Files in CSV on Your Local Computer | 2019-05-30 | 354 | | A Step-by-Step Implementation of Gradient Descent and Backpropagation | 2019-05-30 | 211 | ## Step 2: Convert Month Number to Full Month Name There are several different ways to get a full month name in Pandas and Python. First we will show the built in `dt` method - `.dt.month_name()`: ```python df['month_full'] = df['date'].dt.month_name() ``` This will return the full month name as a new column: | title | date | claps | month\_full | | ------------------------------------------------------------------------------------------------------------------------ | ---------- | ----- | ----------- | | How to Get and Keep Clients as a Freelancer | 2019-05-27 | 965 | May | | A Battle for the Soul of BreadTube Is Currently Taking Place | 2019-07-29 | 105 | July | | Business Opportunity: The Colombian Oils and Lubricants Market | 2019-02-09 | 0 | February | | How to know your price is right | 2019-12-05 | 102 | December | | 3 Types of Reports That Business Analysts Need to Learn | 2019-09-18 | 309 | September | ## Step 3: Convert Month Number to Month abbreviation In this step instead of the full month name we will get an abbreviation of the first 3 letters of the name. It's going to use the same method plus `str` slicing: ```python df['month_short'] = df['date'].dt.month_name().str[:3] ``` The result are short names like: ``` Jul Mar Mar Dec Aug ``` ## Step 4: Convert Month Number to Month name with strftime One more way to convert a month to it's name is by using the method - `strftime`. This can be done by: ```python df['date'].dt.strftime('%b') ``` this will return short names as: ``` Jul Mar Mar Dec Aug ``` For a full month name you can use `%B`: ```python df['date'].dt.strftime('%B') ``` output: ``` July March ``` ## Step 5: Convert Month Number to custom name with mapping Finally, what if you like to get custom labels or month names in different languages like Spanish or French. Then you can apply a custom mapping with `apply`: ```python import pandas as pd month_labels = {1: 'I', 2: 'II', 3: 'III', 4: 'IV', 5: 'V', 6: 'VI', 7: 'VII', 8: 'VIII', 9: 'IX', 10: 'X', 11: 'XI', 12: 'XII'} df['month'] = df['date'].dt.month df['month'].apply(lambda x: month_labels[x]) ``` Above you can find the mapping - `month_labels` which can be updated to suit your needs. You need also to get the month number by `df['date'].dt.month` result is: ``` III VII ``` ## Resources - [Notebook](https://github.com/softhints/datascientyst/blob/master/datetime/3.convert-month-number-to-month-name-pandas-dataframe.ipynb?ref=datascientyst.com) - [pandas.Series.dt.month\_name](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.month%5Fname.html?ref=datascientyst.com) - [pandas.to\_datetime](https://pandas.pydata.org/docs/reference/api/pandas.to%5Fdatetime.html?highlight=to%5Fdatetime&ref=datascientyst.com#pandas.to%5Fdatetime) - [pandas.Series.dt.strftime](https://pandas.pydata.org/docs/reference/api/pandas.Series.dt.strftime.html?ref=datascientyst.com) ### GroupBy and Count Unique Rows in Pandas URL: https://datascientyst.com/pandas-groupby-count/ Last updated: 2022-03-24T07:00:22.000Z In this short guide, we'll see how to use **groupby() on several columns and count unique rows in Pandas**. Several examples will explain how to **group by and apply statistical functions like: `sum`, `count`, `mean`** etc. Often there is a need to group by a column and then get `sum()` and `count()`. × **Pro Tip 1** It's recommended to use method **df.value\_counts** for counting the size of groups in Pandas. It's a bit faster and support parameter \`dropna\` since Pandas 1.3 × **Pro Tip 2** Sorting the results of **groupby/count** or **value\_counts** will slow the process with roughly 20% × **Warning** Be careful of counting NaN values. They can change the expected results and counts. For **value\_counts** use parameter **dropna=True** to count with NaN values. To start, here is the syntax that we may apply in order to combine **`groupby` and `count` in Pandas**: ```python df.groupby(['publication', 'date_m'])['url'].count() ``` The DataFrame used in this article is available from Kaggle. If you like to learn more about how to read Kaggle as a Pandas DataFrame check this article: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) | url | date | publication | | ----------------------------------------------------------------------------------------------------------------- | ---------- | -------------------- | | https://towardsdatascience.com/a-beginners-guide-to-word-embedding-with-gensim-word2vec-model-5970fa56cc92 | 2019-05-30 | Towards Data Science | | https://towardsdatascience.com/hands-on-graph-neural-networks-with-pytorch-pytorch-geometric-359487e221a8 | 2019-05-30 | Towards Data Science | | https://towardsdatascience.com/how-to-use-ggplot2-in-python-74ab8adec129 | 2019-05-30 | Towards Data Science | | https://towardsdatascience.com/databricks-how-to-save-files-in-csv-on-your-local-computer-3d0c70e6a9ab | 2019-05-30 | Towards Data Science | | https://towardsdatascience.com/a-step-by-step-implementation-of-gradient-descent-and-backpropagation-d58bda486110 | 2019-05-30 | Towards Data Science | ## Step 1: Use groupby() and count() in Pandas Let say that we would like to **combine `groupby` and then get unique count per group**. In this first step we will **count the number of unique publications per month from the DataFrame** above. First we need to convert date to month format - YYYY-MM with(learn more about it - [Extract Month and Year from DateTime column in Pandas](https://datascientyst.com/extract-month-and-year-datetime-column-in-pandas/) ```python df['date'] = pd.to_datetime(df['date']) df['date_m'] = df['date'].dt.to_period('M') ``` and then we can **group by two columns - `'publication', 'date_m'` and count the URLs per each group**: ```python df.groupby(['publication', 'date_m'])['url'].count() ``` this will result into: ``` publication date_m Better Humans 2019-03 5 2019-04 4 2019-05 4 2019-06 1 2019-07 3 .. UX Collective 2019-08 24 2019-09 33 2019-10 86 2019-11 28 2019-12 46 ``` An important note is that will **compute the count of each group, excluding missing values**. If we like to [count distinct values in Pandas - nunique()](https://datascientyst.com/ghost/#/editor/post/61821ebfc41e2d0410d3ddf4) \- check the linked article. ## Step 2: groupby(), count() and sum() in Pandas In Pandas method `groupby` will return object which is: `` \- this can be checked by `df.groupby(['publication', 'date_m'])`. This kind of object has an `agg` function which can take a list of aggregation methods. This is very useful if we need to check **multiple statistics methods - `sum(), count(), mean()` per group**. So if we like to group by two columns `publication` and `date_m` \- then to check next aggregation functions - `mean`, `sum`, and `count` we can use: ```python df.groupby(['publication', 'date_m']).agg(['mean', 'count', 'sum']) ``` we will get a MultiIndex as result: | | | reading\_time | | | | ---------------- | --------- | ------------- | ----- | --- | | | | mean | count | sum | | publication | date\_m | | | | | Better Humans | 2019-03 | 12.000000 | 5 | 60 | | 2019-04 | 7.500000 | 4 | 30 | | | 2019-05 | 25.000000 | 4 | 100 | | | 2019-06 | 5.000000 | 1 | 5 | | | 2019-07 | 22.333333 | 3 | 67 | | | 2019-09 | 12.000000 | 3 | 36 | | | 2019-10 | 9.750000 | 4 | 39 | | | 2019-11 | 9.000000 | 2 | 18 | | | 2019-12 | 9.500000 | 2 | 19 | | | Better Marketing | 2019-03 | 9.600000 | 5 | 48 | | 2019-04 | 5.500000 | 4 | 22 | | | 2019-05 | 6.529412 | 34 | 222 | | | 2019-06 | 7.000000 | 15 | 105 | | | 2019-07 | 6.106383 | 47 | 287 | | | 2019-08 | 5.880000 | 25 | 147 | | ## Step 3: groupby() + count() + sort() = value\_counts() In the latest versions of pandas (>= 1.1) you can use **`value_counts` in order to achieve behavior similar to `groupby` and `count`.** ```python df.value_counts(['publication', 'date_m']) ``` will give us: ``` publication date_m The Startup 2019-05 656 2019-10 477 2019-07 401 2019-12 310 2019-06 286 ... Better Humans 2019-07 3 UX Collective 2019-01 2 Better Humans 2019-12 2 2019-11 2 2019-06 1 Length: 79, dtype: int64 ``` Note the code above is equivalent to: ```python df.groupby(['publication', 'date_m'])['url'].count().sort_values(ascending=False) ``` **Note:** Which solution is better depends on the data and the context. In the end of the post there is a performance comparison of both methods. ## Step 4: Combine groupby() and size() Alternative solution is to use **`groupby` and `size` in order to count the elements per group in Pandas**. The example below demonstrate the usage of `size()` \+ `groupby()`: ```python df.groupby(['publication', 'date_m']).size() ``` result is a Pandas series like: ``` publication date_m Better Humans 2019-03 5 2019-04 4 2019-05 4 2019-06 1 2019-07 3 .. UX Collective 2019-08 24 2019-09 33 2019-10 86 2019-11 28 2019-12 46 Length: 79, dtype: int64 ``` ## Step 5: Combine groupby() and describe() The final option is to use the method `describe()`. It will return statistical information which can be extremely useful like: ```python df.groupby(['publication', 'date_m'])['reading_time'].describe() ``` which will return additional stats like: - `count` - `mean` - `std` - percentile etc | | | count | mean | std | min | 25% | 50% | 75% | max | | ---------------- | ------- | --------- | --------- | -------- | ----- | ----- | ----- | ----- | ---- | | publication | date\_m | | | | | | | | | | Better Humans | 2019-03 | 5.0 | 12.000000 | 5.196152 | 3.0 | 13.00 | 13.0 | 15.00 | 16.0 | | 2019-04 | 4.0 | 7.500000 | 5.196152 | 2.0 | 4.25 | 7.0 | 10.25 | 14.0 | | | 2019-05 | 4.0 | 25.000000 | 7.958224 | 17.0 | 21.50 | 23.5 | 27.00 | 36.0 | | | 2019-06 | 1.0 | 5.000000 | NaN | 5.0 | 5.00 | 5.0 | 5.00 | 5.0 | | | 2019-07 | 3.0 | 22.333333 | 15.307950 | 13.0 | 13.50 | 14.0 | 27.00 | 40.0 | | | 2019-09 | 3.0 | 12.000000 | 4.582576 | 8.0 | 9.50 | 11.0 | 14.00 | 17.0 | | | 2019-10 | 4.0 | 9.750000 | 5.123475 | 3.0 | 7.50 | 10.5 | 12.75 | 15.0 | | | 2019-11 | 2.0 | 9.000000 | 1.414214 | 8.0 | 8.50 | 9.0 | 9.50 | 10.0 | | | 2019-12 | 2.0 | 9.500000 | 2.121320 | 8.0 | 8.75 | 9.5 | 10.25 | 11.0 | | | Better Marketing | 2019-03 | 5.0 | 9.600000 | 8.080842 | 5.0 | 6.00 | 6.0 | 7.00 | 24.0 | | 2019-04 | 4.0 | 5.500000 | 1.000000 | 5.0 | 5.00 | 5.0 | 5.50 | 7.0 | | | 2019-05 | 34.0 | 6.529412 | 4.039467 | 2.0 | 5.00 | 5.0 | 7.00 | 20.0 | | | 2019-06 | 15.0 | 7.000000 | 2.725541 | 3.0 | 5.50 | 6.0 | 8.50 | 13.0 | | | 2019-07 | 47.0 | 6.106383 | 2.664044 | 3.0 | 4.00 | 5.0 | 7.00 | 13.0 | | | 2019-08 | 25.0 | 5.880000 | 2.818392 | 3.0 | 4.00 | 5.0 | 6.00 | 15.0 | | ## Performance - groupby() & count() vs value\_counts() Finally lets do a quick comparison of performance between: - `groupby()` and `count()` - `value_counts()` The next example will return equivalent results: ```python df.groupby(['publication', 'date_m'])['publication'].count() ``` ```python df.value_counts(subset=['publication', 'date_m'], sort=False) ``` Below you can find the timings: `df.groupby` ``` 3.59 ms ± 24.1 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) ``` `df.value_counts` ``` 1.05 ms ± 4.73 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) ``` `df.value_counts` with `sort` parameter ``` 1.2 ms ± 11.2 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each) ``` ## Conclusion and Resources In this post we covered how to use **groupby() and count unique rows in Pandas.** How to sort results of `groupby()` and `count()`. Also we covered applying `groupby()` on multiple columns with multiple agg methods like `sum()`, `min()`, `min()`. Finally we saw how to use `value_counts()` in order to count unique values and sort the results. The resources mentioned below will be extremely useful for further analysis: - [Notebook - 1.pandas-groupby-count](https://github.com/softhints/datascientyst/blob/master/groupby/1.pandas-groupby-count.ipynb?ref=datascientyst.com) - [pandas.DataFrame.groupby](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.groupby.html?highlight=groupby&ref=datascientyst.com) - [pandas.DataFrame.agg](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.agg.html?ref=datascientyst.com) - [GroupBy.describe](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.core.groupby.DataFrameGroupBy.describe.html?ref=datascientyst.com) - [GroupBy.count](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.core.groupby.DataFrameGroupBy.count.html?ref=datascientyst.com) - [pandas.DataFrame.value\_counts](https://pandas.pydata.org/pandas-docs/dev/reference/api/pandas.DataFrame.value%5Fcounts.html?ref=datascientyst.com) ### How to rename column in Pandas URL: https://datascientyst.com/pandas-rename-column-names/ Last updated: 2023-01-10T09:21:36.000Z In this short guide, I'll show you how to **rename column names in Pandas DataFrame**. **(1) rename single column** ```python df.rename(columns = {'$b':'B'}, inplace = True) ``` **(2) rename multiple columns** ```python column_map = {'A': 'First', 'B': 'Second'} df = df.rename(columns=column_map) ``` **(3) rename multi-index columns** ```python cols = pd.MultiIndex.from_tuples([(0, 1), (0, 2)]) df = pd.DataFrame([[1,2], [3,4]], columns=cols) ``` **(4) rename all columns** ```python df.columns = ['a', 'b'] ``` In the next sections, I'll review the steps to apply the above syntax in practice and a few exceptional cases. Let's say that you have the following DataFrame with random numbers generated by: ```python import pandas as pd import numpy as np df = pd.DataFrame(np.random.randint(0,10,size=(5, 5)), columns=list('ABCDF')) ``` DataFrame: | | A | B | C | D | F | | - | - | - | - | - | - | | 0 | 4 | 8 | 9 | 0 | 1 | | 1 | 5 | 8 | 6 | 1 | 0 | | 2 | 7 | 9 | 8 | 1 | 1 | | 3 | 6 | 8 | 3 | 8 | 9 | | 4 | 6 | 0 | 2 | 8 | 8 | If you like to understand more about how to create DataFrame with random numbers please check: [How to Create a Pandas DataFrame of Random Integers](https://datascientyst.com/how-to-create-a-dataframe-of-random-integers-with-pandas/) ## Step 1: Rename all column names in Pandas DataFrame Column names in Pandas DataFrame can be accessed by attribute: `.columns`: ```python df.columns ``` result: ``` Index(['A', 'B', 'C', 'D', 'F'], dtype='object') ``` The same attribute can be used to **rename all columns in Pandas.** ```python df.columns = ['First', 'Second', '3rd', '4th', '5th'] ``` ```python df.columns ``` result: ``` Index(['First', 'Second', '3rd', '4th', '5th'], dtype='object') ``` × **Pro Tip** To **rename single column in Pandas** use: df.rename(columns = {'$b':'B'}, inplace = True) ## Step 2: Rename specific column names in Pandas If you like to rename specific columns in Pandas you can use method - `.rename`. Let's work with the first DataFrame with names - A, B etc. To rename two columns - `A, B` to `First, Second` we can use the following code: ```python column_map = {'A': 'First', 'B': 'Second'} df = df.rename(columns=column_map) ``` will result in: ``` Index(['First', 'Second', 'C', 'D', 'F'], dtype='object') ``` **Note**: If any of the column names are missing they will be skipped without any error or warning because of default parameter `errors='ignore'` **Note 2:** Instead of syntax: `df = df.rename(columns=column_map)` you can use `df.rename(columns=column_map, inplace=False)` ![pandas-rename-column-names](https://datascientyst.com/content/images/2021/08/pandas-rename-column-names.png) × **Pro Tip** Method **rename** can take a function: df.rename(columns=lambda x: x.lstrip()) × **Warning** Don't forget to use **inplace=True** to make the renaming permanent! ## Step 3: Rename column names in Pandas with lambda Sometimes you may like to replace a character or apply other functions to DataFrame columns. In this example we will **change all columns names from upper to lowercase**: ```python df = df.rename(columns=lambda x: x.lower()) ``` the result will be: ``` Index(['a', 'b', 'c', 'd', 'f'], dtype='object') ``` This step is suitable for complex transformations and logic. ## Step 4: Rename column names in Pandas with str methods You can apply str methods to Pandas columns. For example we can add extra character for each column name with a regex: ```python df.columns = df.columns.str.replace(r'(.*)', r'Column \1') ``` Working with the original DataFrame will give us: ``` Index(['Column A', 'Column B', 'Column C', 'Column D', 'Column F'], dtype='object') ``` ## Step 5: Rename multi-level column names in DataFrame Finally let's check how to rename columns when you have MultiIndex. Let's have a DataFrame like: ```python import pandas as pd cols = pd.MultiIndex.from_tuples([(0, 1), (0, 2)]) df = pd.DataFrame([[1,2], [3,4]], columns=cols) ``` | | 0 | | | - | - | - | | | 1 | 2 | | 0 | 1 | 2 | | 1 | 3 | 4 | If we check the column names we will get: ```python df.columns ``` ``` MultiIndex([(0, 1), (0, 1)], ) ``` Renaming of the MultiIndex columns can be done by: ```python df.columns = pd.MultiIndex.from_tuples([('A', 'B'), ('A', 'C')]) ``` and will result into: | | A | | | - | - | - | | | B | C | | 0 | 1 | 2 | | 1 | 3 | 4 | ## Resources - [Notebook - pandas-rename-column-names.ipynb](https://github.com/softhints/datascientyst/blob/master/column/2.pandas-rename-column-names.ipynb?ref=datascientyst.com) - [pandas.DataFrame.rename — pandas](https://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.rename.html?ref=datascientyst.com) - [pandas.Series.str.replace](https://pandas.pydata.org/docs/reference/api/pandas.Series.str.replace.html?ref=datascientyst.com) - [pandas.DataFrame.columns](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.columns.html?ref=datascientyst.com) ### How to Convert DateTime to Quarter in Pandas URL: https://datascientyst.com/convert-datetime-to-quarter-in-pandas/ Last updated: 2022-10-28T06:08:39.000Z Need to rename convert string or datetime to quarter in Pandas? If so, you may use the following syntax to extract quarter info from your DataFrame: ```python df['Date'].dt.to_period('Q') ``` The picture below demonstrate it: ![convert-datetime-to-quarter-in-pandas](https://datascientyst.com/content/images/2022/10/convert-datetime-to-quarter-in-pandas.webp) To begin, let's create a simple DataFrame with a datetime column: ```python import pandas as pd dates = ['2021-01-01', '2021-05-02', '2021-08-03'] df = pd.DataFrame({'StartDate': dates}) df['StartDate'] = pd.to_datetime(df['StartDate']) df ``` | StartDate | | ---------- | | 2021-01-01 | | 2021-05-02 | | 2021-08-03 | ## Step 1: Extract Quarter as YYYYMM from DataTime Column in Pandas Let's say that you want to extract the quarter info in the next format: `YYYYMM`. Pandas offers a built-in method for this purpose `.dt.to_period('Q')`: ```python df['quarter'] = df['StartDate'].dt.to_period('Q') ``` the result would be: ``` 0 2021Q1 1 2021Q2 2 2021Q3 ``` ## Step 2: Extract Quarter(custom format) from DataTime Column in Pandas What if you like to get a custom format instead of the default one? In this case you can use method: `dt.strftime('q%q %y\'')` to change the format like: ```python df['quarter_custom'] = df['StartDate'].dt.to_period('Q').dt.strftime('q%q %y\'') ``` This will output: ``` 0 q1 21' 1 q2 21' 2 q3 21' ``` Another option for custom formatting might be separated extraction of the information as: ```python df['StartDate'].dt.year.astype(str) + '-' + df['StartDate'].dt.quarter.astype(str) ``` ## Step 3: Extract Quarters without the year If you need to analyse the quarter info without the year information you can get it by: ```python df['StartDate'].dt.quarter ``` ## Resources - [Notebook - Convert DateTime to Quarter in Pandas](https://github.com/softhints/datascientyst/blob/master/datetime/2.convert-datetime-to-quarter-in-pandas.ipynb?ref=datascientyst.com) ### How to style boolean values by different colors in Pandas URL: https://datascientyst.com/style-boolean-values-different-colors-in-pandas/ Last updated: 2021-08-14T08:58:35.000Z In this article, you can find two approaches to color boolean columns in a Pandas DataFrame. The first one will change the background color while the second one will change the font color. ## Step 1: Create Pandas DataFrame with boolean values To start, let's create a DataFrame with several boolean columns which we will color. For this article we are going to read data from Kaggle and search for a list of keywords - more info here: [How to Add New Column Based on List of Keywords in Pandas DataFrame](https://datascientyst.com/add-new-column-list-keywords-pandas-dataframe/) You can download the data and read it manually by: ```python import pandas as pd df = pd.read_csv('../data/medium_data.csv.zip') ``` in order to create multiple boolean columns we will search for list of words in a given column: ```python keywords=['how', 'data', 'question', 'guide'] df['keyword'] = df['title'].str.findall('|'.join(keywords)).apply(set).str.join(', ') df[df['keyword'].str.len() > 5][['url', 'title', 'keyword']] ``` So this will be our starting data: | title | keyword | how | guide | data | question | | ------------------------------------------------------------------------------------------------------------------- | --------- | ----- | ----- | ----- | -------- | | Training on batch: how to split data effectively? | how, data | True | False | True | False | | Design & Data: how to humanize data | how, data | True | False | True | False | | Fyre Festival achieved perfect product-market fit (and that’s why we should question the Lean Startup and VC dogma) | question | False | False | False | True | | Using a ‘sneak attack’ question during your Designer interviews | question | False | False | False | True | | A question you may never have asked — culture fit or culture add? | question | False | False | False | True | ## Step 2: Color background color of boolean column in Pandas First we will demonstrate how to **change the background color in a single or multiple boolean columns in Pandas DataFrame**. For this purpose we will define a function which will add red background for `True` and green for `False`: ```python def color_boolean(val): color ='' if val == True: color = 'red' elif val == False: color = 'green' return 'background-color: %s' % color ``` The function can be used for a **single or several columns** like this: ```python df.style.applymap(color_boolean, subset=['data']) ``` and for **all boolean columns in a DataFrame** like this: ```python df.style.applymap(color_boolean) ``` The result is: ![style-boolean-values-different-colors-in-pandas](https://datascientyst.com/content/images/2021/08/style-boolean-values-different-colors-in-pandas.png) ## Step 3: Change font color of boolean values in Pandas Next if you like to highlight the `True` and `False` values of Pandas DataFrame then you can change only the font color. Let's define a different function: ```python def color_boolean(val): color ='' if val == True: color = 'red' elif val == False: color = 'green' return 'color: %s' % color ``` It can be used in the same way only the result is different: ```python df.style.applymap(color_boolean) ``` ## Step 4: Use heatmap to color boolean values in Pandas The last option which will be demonstrated for styling boolean columns in Pandas is with heatmap technique. First we may need to prepare our data in order to be correctly styled - convert all boolean columns to `int` \- `True` will become 1 and `False` will become 0. After that we can create a simple heatmap by: ```python df_temp[['how', 'guide', 'data', 'question']] = df_temp[['how', 'guide', 'data', 'question']].astype(int) df_temp.style.background_gradient(cmap='Blues') ``` And the result will be: ![style-boolean-values-different-colors-in-pandas-heatmap](https://datascientyst.com/content/images/2021/08/style-boolean-values-different-colors-in-pandas-heatmap.png) ## Notebook - [Style Boolean Values by Different Colors in Pandas](https://github.com/softhints/datascientyst/blob/master/styling/2.style-boolean-values-different-colors-in-pandas.ipynb?ref=datascientyst.com) ### How to Add New Column Based on List of Keywords in Pandas DataFrame URL: https://datascientyst.com/add-new-column-list-keywords-pandas-dataframe/ Last updated: 2021-12-02T10:56:34.000Z In this short guide, I'll show you the steps **to extract a list of keywords matched in a Pandas DataFrame and create new column(s)**. In particular, I'll show you how to return keywords from a given column. The image below shows what is the final outcome: ![add-new-column-list-keywords-pandas-dataframe](https://datascientyst.com/content/images/2021/08/add-new-column-list-keywords-pandas-dataframe.png) To start with a simple example, let's say that you have the next Pandas DataFrame: | url | title | keyword | | -------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | --------- | | https://towardsdatascience.com/training-on-batch-how-to-split-data-effectively-3234f3918b07 | Training on batch: how to split data effectively? | how, data | | https://uxdesign.cc/design-and-data-how-to-humanize-data-32a03079311f | Design & Data: how to humanize data | how, data | | https://medium.com/swlh/fyre-festival-achieved-perfect-product-market-fit-and-thats-why-we-should-question-the-lean-a6a45fcb735a | Fyre Festival achieved perfect product-market fit (and that's why we should question the Lean Startup and VC dogma) | question | | https://uxdesign.cc/using-a-sneak-attack-question-during-your-designer-interviews-b918ff600977 | Using a ‘sneak attack’ question during your Designer interviews | question | | https://medium.com/swlh/a-question-you-may-never-have-asked-culture-fit-or-culture-add-cf65b00770cf | A question you may never have asked — culture fit or culture add? | question | Notebook with the code: [Extract list of keywords from a column in Pandas](https://github.com/softhints/datascientyst/blob/master/column/1.add-new-column-list-keywords-pandas-dataframe.ipynb?ref=datascientyst.com) ## Step 1: Read test DataFrame from Kaggle The DataFrame above is available from Kaggle. If you like to learn more about how to read Kaggle as a Pandas DataFrame check this article: [How to Search and Download Kaggle Dataset to Pandas DataFrame](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) For this article we will use next code to download and read it: ```python import kaggle kaggle.api.authenticate() kaggle.api.dataset_download_file('dorianlazar/medium-articles-dataset', file_name='medium_data.csv', path='data/') ``` read it by: ```python import pandas as pd df = pd.read_csv('data/medium_data.csv.zip') ``` ## Step 2: Extract list of keywords from a column to new column At this point we will define **the list of keywords which we like to extract**: ```python keywords = ['how', 'data', 'question', 'guide'] ``` Then we are going to **perform the extraction and the addition to a new column**: ```python df['keyword'] = df['title'].str.findall('|'.join(keywords)).apply(set).str.join(', ') ``` At this point the new column `keyword` will contain all keywords found separated by a comma. ## Step 3: Extract list of keywords to multiple columns To extract the list of keywords to different columns use the next syntax: ```python for keyword in keywords: df[keyword] = df['title'].str.contains(keyword) ``` This will iterate over the list of the all columns and create a new column with `True` or `False` if the word exists or not. If you like to style your DataFrame in the same way please check: [How to style boolean values by different colors in Pandas](https://datascientyst.com/style-boolean-values-different-colors-in-pandas/) Final result can be found on the image below: ![extract-list-keywords-new-columns-pandas-dataframe](https://datascientyst.com/content/images/2021/08/extract-list-keywords-new-columns-pandas-dataframe.png) ### How to Extract Month and Year from DateTime column in Pandas URL: https://datascientyst.com/extract-month-and-year-datetime-column-in-pandas/ Last updated: 2022-10-28T06:20:48.000Z In this short guide, I'll show you how to extract Month and Year from a DateTime column in Pandas DataFrame. You can also find how to convert string data to a DateTime. So at the end you will get: `01/08/2021` \-> `2021-08` `DD/MM/YYYY` \-> `YYYY-MM` or any other date format. We will also cover `MM/YYYY`. To start, here is the syntax that you may apply in order extract concatenation of year and month: ```python .dt.to_period('M') ``` In the next section, I'll review the steps to apply the above syntax in practice. ## Step 1: Create a DataFrame with Datetime values Lets create a DataFrame which has a single column StartDate: ```python dates = ['2021-08-01', '2021-08-02', '2021-08-03'] df = pd.DataFrame({'StartDate': dates}) ``` result: | StartDate | | ---------- | | 2021-08-01 | | 2021-08-02 | | 2021-08-03 | In order to convert string to Datetime column we are going to use: ```python df['StartDate'] = pd.to_datetime(df['StartDate']) ``` ## Step 2: Extract Year and Month with .dt.to\_period('M') - format YYYY-MM In order to extract from a full date only the year plus the month: 2021-08-01 -> 2021-08 we need just this line: ```python df['StartDate'].dt.to_period('M') ``` result: ``` 0 2021-08 1 2021-08 2 2021-08 ``` ## Step 3: Extract Year and Month other formats MM/YYYY What if you like to get the month first and then the year? In this case we will use `.dt.strftime` in order to produce a column with format: `MM/YYYY` or any other format. ```python df['StartDate'].dt.strftime('%m/%Y') ``` ``` 0 08/2021 1 08/2021 2 08/2021 ``` Note: have in mind that this solution might be really slow in case of a huge DataFrame. ## Step 4: Extracting Year and Month separately and combine them A bit faster solution than step 3 plus a trace of the month and year info will be: - extract month and date to separate columns - combine both columns into a single one ```python df['yyyy'] = pd.to_datetime(df['StartDate']).dt.year df['mm'] = pd.to_datetime(df['StartDate']).dt.month ``` | StartDate | yyyy | mm | | ---------- | ---- | -- | | 2021-08-01 | 2021 | 8 | | 2021-08-02 | 2021 | 8 | | 2021-08-03 | 2021 | 8 | and then: ```python df['yyyy'].astype(str) + '-'+ df['mm'].astype(str) ``` Note: If you don't need extra columns you can just do: ```python df['StartDate'].dt.year.astype(str) + "-" + df['StartDate'].dt.month.astype(str) ``` Notebook with all examples: [Extract Month and Year from DateTime column](https://github.com/softhints/datascientyst/blob/master/datetime/1.extract-month-and-year-datetime-column-in-pandas.ipynb?ref=datascientyst.com) ![extract-month-year-datetime-column-pandas](https://datascientyst.com/content/images/2022/10/extract-month-year-datetime-column-pandas.webp) ### Read Excel XLS with Python Pandas URL: https://datascientyst.com/read-excel-xls-with-python-pandas/ Last updated: 2021-07-18T09:47:56.000Z In this post you can learn **how to read Excel files (ext xls, xlsx etc) with Python and Pandas**. We will import one or several sheets from an Excel file to a Pandas DataFrame. The list of the supported file extensions: - `xls` - `xlsx` - `xlsm` - `xlsb` - `odf` - `ods` - `odt` Note for `ods`, `ods` and `odt` please check: [Read Excel(OpenDocument ODS) with Python Pandas](https://datascientyst.com/read-excel-opendocument-ods-python-pandas) ## Step 1: Install Pandas and odfpy Python offers many different modules for reading and manipulating Excel files. In this guide we are going to use `pandas` and `odfpy`: ```python pip install pandas pip install odfpy ``` ## Step 2: Read the one sheet of Excel(XLS) file Pandas offers a powerful method for **reading any type of Excel files `read_excel()`**. It's pretty easy to be used and requires only the file path: ```python import pandas as pd pd.read_excel('animals.xls') ``` It will read and return all non empty cells from the Excel file: | | Rank | Animal | Maximum speed | Class | Notes | | - | ---- | ------------------------------- | ------------------------------------------------- | ------------- | ------------------------------------------------- | | 0 | 1 | Peregrine falcon | 389 km/h (242 mph)108 m/s (354 ft/s)\[2\]\[6\] | Flight-diving | The peregrine falcon is the fastest aerial ani... | | 1 | 2 | Golden eagle | 240–320 km/h (150–200 mph)67–89 m/s (220–293 f... | Flight-diving | Assuming the maximum size at 1.02 m, its relat... | | 2 | 3 | White-throated needletail swift | 169 km/h (105 mph)\[8\]\[9\]\[10\] | Flight | NaN | | 3 | 4 | Eurasian hobby | 160 km/h (100 mph)\[11\] | Flight | Can sometimes outfly the swift | | 4 | 5 | Mexican free-tailed bat | 160 km/h (100 mph)\[12\] | Flight | It has been claimed to have the fastest horizo... | | 5 | 6 | Frigatebird | 153 km/h (95 mph) | Flight | The frigatebird's high speed is helped by its ... | | 6 | 7 | Rock dove (pigeon) | 148.9 km/h (92.5 mph)\[13\] | Flight | Pigeons have been clocked flying 92.5 mph (148... | | 7 | 8 | Spur-winged goose | 142 km/h (88 mph)\[14\] | Flight | NaN | | 8 | 9 | Gyrfalcon | 128 km/h (80 mph)\[citation needed\] | Flight | NaN | ![pandas-read excel-xls-xlsx-python](https://datascientyst.com/content/images/2021/07/pandas-read-excel-xls-xlsx-python.png) ## Step 3: Read the second sheet of Excel file by name If you like to **read data from a specific sheet** \- for example `Sheet 2` then you can specify the name as a parameter - `sheet_name`: ```python pd.read_excel('animals.xlsx', sheet_name="Sheet2") ``` Which will result in: | | Blackbuck | Unnamed: 1 | | - | ------------------------------------------------------- | ------------------------------------------------- | | 0 | NaN | NaN | | 1 | Male blackbuck | Male blackbuck | | 2 | NaN | NaN | | 3 | Female with young at the National Zoological Park Delhi | Female with young at the National Zoological P... | | 4 | Conservation status | Conservation status | | 5 | Least Concern (IUCN 3.1)\[1\] | Least Concern (IUCN 3.1)\[1\] | | 6 | Scientific classification | Scientific classification | ## Step 4: Python read excel file - specify columns and rows If you like to read a range of data and not the whole sheet - `read_excel` offers several very useful parameters. ### Python read excel file select rows Next code example will show you **how to read 3 rows skipping the first two rows**. In this way Pandas will read only some rows from the whole sheet: ```python pd.read_excel('animals.xlsx', skiprows=2, nrows=3) ``` which will result in: | | 2 | Golden eagle | 240–320 km/h (150–200 mph)67–89 m/s (220–293 f... | Flight-diving | Assuming the maximum size at 1.02 m, its relat... | | - | - | ------------------------------- | ------------------------------------------------- | ------------- | ------------------------------------------------- | | 0 | 3 | White-throated needletail swift | 169 km/h (105 mph)\[8\]\[9\]\[10\] | Flight | NaN | | 1 | 4 | Eurasian hobby | 160 km/h (100 mph)\[11\] | Flight | Can sometimes outfly the swift | | 2 | 5 | Mexican free-tailed bat | 160 km/h (100 mph)\[12\] | Flight | It has been claimed to have the fastest horizo... | ### Python read excel file select columns If you like to\*\* work with few columns\*\* and not the whole sheet - then parameter `use_cols` can be used as shown: ```python pd.read_excel('animals.xlsx', usecols='C:D') ``` ### Python read excel file specify columns and rows Finally if you like to **select a range from specific columns and rows** than you can use: ```python ``` Which will result into: | | 240–320 km/h (150–200 mph)67–89 m/s (220–293 f... | Flight-diving | | - | ------------------------------------------------- | ------------- | | 0 | 169 km/h (105 mph)\[8\]\[9\]\[10\] | Flight | | 1 | 160 km/h (100 mph)\[11\] | Flight | | 2 | 160 km/h (100 mph)\[12\] | Flight | ## Step 5\. Read multiple sheets from Excel file What if you like to r**ead with Pandas multiple sheets from Excel.** It's possible with `pd.read_excel` by providing a list of all sheets to be read as follows: ```python pd.read_excel('animals.xlsx', sheet_name=["Sheet1", "Sheet2"]) ``` Note that a dictionary of - keys - sheet names - values - resulted DataFrames will be returned. In order to access data you can access it by a sheet name as: ```python pd.read_excel('animals.xlsx', sheet_name=["Sheet1", "Sheet2"]).get('Sheet1') ``` which will return the data for `Sheet1` as a DataFrame. ### Read All Sheets For loading all sheets from Excel file use `sheet_name=None`: ```python pd.read_excel('animals.xlsx', sheet_name=None) ``` ## Step 6\. Pandas read excel data with conversion, NA values and parsing Finally let's check what we can do if we need to convert data, drop or fill missing values, parse dates and numbers. Pandas offers several parameters for this purpose: - **converters** \- dict of functions for converting values in certain columns - **keep\_default\_na** \- whether or not to include the default NaN values - **parse\_dates** - **ate\_parser** \- converting a sequence of string columns to an array of datetime instances. - **thousands** - **convert\_float** You can check the Notebook in the resources for more examples of the above. ## Resources - [Python Pandas Reading Excel files](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/io.html?ref=datascientyst.com#excelfile-class) - [pandas.read\_excel](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read%5Fexcel.html?ref=datascientyst.com#pandas.read%5Fexcel) - [Notebook - Read Excel ODS with Python Pandas](https://github.com/softhints/datascientyst/blob/master/excel/1.read-excel-xls-xlsx-python-pandas.ipynb?ref=datascientyst.com) ### Read Excel(OpenDocument ODS) with Python Pandas URL: https://datascientyst.com/read-excel-opendocument-ods-python-pandas/ Last updated: 2022-05-13T07:33:21.000Z In this post we will see **how to read Excel files with extensions:ods and . fods with Python and Pandas**. Pandas offers a read\_excel() method to read Excel files as a DataFrame. Many options are available - you can read any sheet, all sheets, first one or range of data. Pandas read and convert any data stored in ODS format to the popular DataFrame format. Let's assume the next ODS file stored on our local machine which needs to be read as a DataFrame or any othder Python format: ![read-excel-opendocument-ods-python-pandas](https://datascientyst.com/content/images/2021/07/read-excel-opendocument-ods-python-pandas.png) ## Step 1: Install Pandas and odfpy Before reading the ODS files with Python we need to install additional packages like: `odfpy` \+ Pandas (if not installed or upgrade to latest). To install `odfpy` \+ Pandas use next commands: ```python pip install pandas pip install odfpy ``` ## Step 2: Read the first sheet of Excel(ODS) files Now we can read the first sheet of the worksheet by calling Pandas method `read_excel()`: ```python import pandas as pd pd.read_excel('~/Desktop/animals.ods', engine='odf') ``` Which will result in: | | Rank | Animal | Maximum speed | Class | Notes | | - | ---- | ------------------------------- | ------------------------------------------------- | ------------- | ------------------------------------------------- | | 0 | 1 | Peregrine falcon | 389 km/h (242 mph)108 m/s (354 ft/s)\[2\]\[6\] | Flight-diving | The peregrine falcon is the fastest aerial ani... | | 1 | 2 | Golden eagle | 240–320 km/h (150–200 mph)67–89 m/s (220–293 f... | Flight-diving | Assuming the maximum size at 1.02 m, its relat... | | 2 | 3 | White-throated needletail swift | 169 km/h (105 mph)\[8\]\[9\]\[10\] | Flight | NaN | | 3 | 4 | Eurasian hobby | 160 km/h (100 mph)\[11\] | Flight | Can sometimes outfly the swift | | 4 | 5 | Mexican free-tailed bat | 160 km/h (100 mph)\[12\] | Flight | It has been claimed to have the fastest horizo... | | 5 | 6 | Frigatebird | 153 km/h (95 mph) | Flight | The frigatebird's high speed is helped by its ... | | 6 | 7 | Rock dove (pigeon) | 148.9 km/h (92.5 mph)\[13\] | Flight | Pigeons have been clocked flying 92.5 mph (148... | | 7 | 8 | Spur-winged goose | 142 km/h (88 mph)\[14\] | Flight | NaN | | 8 | 9 | Gyrfalcon | 128 km/h (80 mph)\[citation needed\] | Flight | NaN | ## Step 3: Pandas read excel sheet by name To work with multiple sheets `read_excel` method request the path the the file and the sheet name: ```python pd.read_excel("animals.ods", sheet_name="Sheet1") ``` If the sheet name is missed then the first sheet will be read. ## Step 4: Pandas read excel range of data If you like to read only a specific range of data read\_excel has several useful parameters for this purpose like: - **skiprows** \- line numbers to skip. If you like to start from row 6 than use - `skiprows=6` - **nrows** \- how many rows to read. If you like to start with fewer rows for a test - **index\_col** \- which column to be used as a index - **usecols** \- which columns to be read only a specific range of columns: ```python pd.read_excel('/home/vanx/Desktop/animals.ods', usecols="A:C") ``` To find more examples on how to read and import Excel files to Pandas DataFrame and Python check: [Notebook - Read Excel ODS with Python Pandas](https://github.com/softhints/datascientyst/blob/master/excel/2.read-excel-opendocument-ods-python-pandas.ipynb?ref=datascientyst.com) ## Step 5: Pandas read excel - XLRDError: Openoffice.org ODS file; not supported If you face error during reading the ODS or any other excel file with Python and Pandas: > XLRDError: Openoffice.org ODS file; not supported Then you can update Pandas to the latest possible version and this will resolve the problem. Pandas can be upgraded by: ```python pip install --upgrade pandas ``` In some cases you can specify the reading engine for `pd.read_excel` as shown in the example: ```python pd.read_excel('animals.ods', engine='odf') ``` ## Resources - [Python Pandas Reading Excel files](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/io.html?ref=datascientyst.com#excelfile-class) - [pandas.read\_excel](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read%5Fexcel.html?ref=datascientyst.com#pandas.read%5Fexcel) - [Notebook - Read Excel ODS with Python Pandas](https://github.com/softhints/datascientyst/blob/master/excel/2.read-excel-opendocument-ods-python-pandas.ipynb?ref=datascientyst.com) ### How to Get a List of N Different Colors and Names in Python/Pandas URL: https://datascientyst.com/get-list-of-n-different-colors-names-python-pandas/ Last updated: 2021-12-02T10:57:24.000Z Looking to generate a **list of different colors or get color names in Python?** We are going to demonstrate combination of different modules like: - pandas - searborn - webcolors - matplotlib In order to generate, **list color names and values. Also we can see how to work with color palettes and HTML/CSS values.** If so, you'll see the options to accomplish these goals using simple examples. ## Step 1: Generate N Random Colors with Python In this step we will **get a list of many different colors as hex values in Python.** We are going to use pure Python code and will generate a list of unique colors with `randomint` and `for` loop: ```python from random import randint color = [] n = 10 for i in range(n): color.append('#%06X' % randint(0, 0xFFFFFF)) ``` This will generate a list like: > \['#4E8CA1', '#F9C1E7', '#933D05', '#595697', '#417D22', '#7D8377', '#624F7B', '#C25D39', '#A24AFD', '#2AED9E'\] You need to provide just the number of the random colors to be generated. ## Step 2: Display N random colors with Python and JupyterLab **To display the different colors generated from the above code we are going to use Pandas.** First we are going to define a function which is going to display each row of a Pandas DataFrame in a different color(more example on formatting - [How to Set Pandas DataFrame Background Color Based On Condition](https://datascientyst.com/pandas-dataframe-background-color-based-condition-value-alternate-row-color-based-group/)) : ```python from random import randint def format_color_groups(df, color): x = df.copy() i = 0 for factor in color: style = f'background-color: {color[i]}' x.loc[i] = style i = i + 1 return x ``` Next we are going to generate a DataFrame with a random integers (more info - [How to Create a Pandas DataFrame of Random Integers](https://datascientyst.com/how-to-create-a-dataframe-of-random-integers-with-pandas/): ```python from random import randint import numpy as np import pandas as pd color = [] n = 10 df = pd.DataFrame(np.random.randint(0,n,size=(n, 2)), columns=list('AB')) ``` Finally we are going to apply the function to the generated DataFrame and pass the list of colors: ```python df.style.apply(format_color_groups, color=color, axis=None) ``` This will result in: | | A | B | | - | - | - | | 0 | 4 | 4 | | 1 | 3 | 4 | | 2 | 5 | 0 | | 3 | 5 | 1 | | 4 | 8 | 2 | | 5 | 7 | 4 | | 6 | 3 | 2 | | 7 | 5 | 0 | | 8 | 2 | 1 | | 9 | 8 | 9 | ## Step 3: Get color palette with seaborn and Python Python package `seaborn` offers color palettes with different styles. You can find more about it: [seaborn.color\_palette](https://seaborn.pydata.org/generated/seaborn.color%5Fpalette.html?ref=datascientyst.com) Below you can find a simple example: ```python import seaborn as sns sns.color_palette() ``` or ```python sns.color_palette("flare") sns.color_palette("pastel") ``` You can find the generated colors on the image below: ![get-list-of-n-different-colors-names-python](https://datascientyst.com/content/images/2021/07/get-list-of-n-different-colors-names-python.png) How many colors to generate depends on a parameter: ```python palette = sns.color_palette(None, 3) ``` ## Step 4: Get color palette with matplotlib Another option to get different color palettes in Python is with the package matplotlib. After installing it you can generate different color ranges with code like: ```python import numpy as np import matplotlib as mpl import matplotlib.pyplot as plt from matplotlib import cm from collections import OrderedDict cmaps = OrderedDict() cmaps['Miscellaneous'] = [ 'flag', 'prism', 'ocean', 'gist_earth', 'terrain', 'gist_stern', 'gnuplot', 'gnuplot2', 'CMRmap', 'cubehelix', 'brg', 'gist_rainbow', 'rainbow', 'jet', 'turbo', 'nipy_spectral', 'gist_ncar'] gradient = np.linspace(0, 1, 256) gradient = np.vstack((gradient, gradient)) def plot_color_gradients(cmap_category, cmap_list): # Create figure and adjust figure height to number of colormaps nrows = len(cmap_list) figh = 0.35 + 0.15 + (nrows + (nrows - 1) * 0.1) * 0.22 fig, axs = plt.subplots(nrows=nrows + 1, figsize=(6.4, figh)) fig.subplots_adjust(top=1 - 0.35 / figh, bottom=0.15 / figh, left=0.2, right=0.99) axs[0].set_title(cmap_category + ' colormaps', fontsize=14) for ax, name in zip(axs, cmap_list): ax.imshow(gradient, aspect='auto', cmap=plt.get_cmap(name)) ax.text(-0.01, 0.5, name, va='center', ha='right', fontsize=10, transform=ax.transAxes) # Turn off *all* ticks & spines, not just the ones with colormaps. for ax in axs: ax.set_axis_off() for cmap_category, cmap_list in cmaps.items(): plot_color_gradients(cmap_category, cmap_list) plt.show() ``` More information can be found: [Choosing Colormaps in Matplotlib](https://matplotlib.org/stable/tutorials/colors/colormaps.html?ref=datascientyst.com) ![get-list-of-n-different-colors-names-python-mathplotlib](https://datascientyst.com/content/images/2021/07/get-list-of-n-different-colors-names-python-mathplotlib.png) ## Step 5: Working with color names and color values format (HTML and CSS) in Python Additional library `webcolors` is available if you need to: - **get the color name in Python** - **convert color name to HTML/CSS or hex format** More about this Python module: [webcolors](https://pypi.org/project/webcolors/?ref=datascientyst.com). And example usage of converting hex value to color name: ```python import webcolors webcolors.hex_to_name(u'#daa520') webcolors.hex_to_name(u'#000000') webcolors.hex_to_name(u'#FFFFFF') ``` Which will produce: - 'goldenrod' - 'black' - 'white' On the other hand if you like to get the color value based on its name you can do: ```python webcolors.html5_parse_legacy_color(u'lime') webcolors.html5_parse_legacy_color(u'salmon') ``` Will result in: - HTML5SimpleColor(red=0, green=255, blue=0) - HTML5SimpleColor(red=250, green=128, blue=114) ### How to Drop a Level from a MultiIndex in Pandas DataFrame URL: https://datascientyst.com/pandas-drop-multiindex-level/ Last updated: 2021-11-12T12:44:34.000Z Here are several approaches to **drop levels of MultiIndex in a Pandas DataFrame**: - `droplevel` \- completely drop MultiIndex level - `reset_index` \- remove levels of MultiIndex while storing data into columns/rows If you want to find more about: [What is a DataFrame MultiIndex in Pandas](https://datascientyst.com/dataframe-multiindex-in-pandas) ## Step 1: Pandas drop MultiIndex by method - droplevel ### Pandas drop MultiIndex on index/rows Method `droplevel()` will remove one, several or all levels from a MultiIndex. Let's check the default execution by next example: ```python import pandas as pd cols = pd.MultiIndex.from_tuples([(0, 1), (0, 1)]) df = pd.DataFrame([[1,2], [3,4]], index=cols) df ``` | | | 0 | 1 | | - | - | - | - | | 0 | 1 | 1 | 2 | | 1 | 3 | 4 | | We can drop level 0 of the MultiIndex by: ```python df.droplevel(level=0) ``` which will result in: | | 0 | 1 | | - | - | - | | 1 | 1 | 2 | | 1 | 3 | 4 | or level 1: ```python df.droplevel(level=1) ``` ### Pandas drop MultiIndex on columns If the hierarchical indexing is on the columns then we can drop levels by parameter `axis`: ```python df.droplevel(level=0, axis=1) ``` ### Get column names of the dropped columns If you like to get the names of the columns which will be dropped you can use next syntax: - for columns: ```python df.columns.droplevel(level=0) ``` - for rows: ```python df.index.droplevel(level=0) ``` which will result of single index: > Int64Index(\[1, 1\], dtype='int64') ### Pandas drop MultiIndex with `.columns.droplevel(level=0)` So we can drop level of MultiIndex with a simple reset of the column names like: ```python df.columns = df.columns.droplevel(level=0) ``` ![drop-level-from-multiindex-pandas](https://datascientyst.com/content/images/2021/07/drop-level-from-multiindex-pandas.png) ## Step 2: Pandas drop MultiIndex to column values by reset\_index ### Drop all levels of MultiIndex to columns Use reset\_index if you like to drop the MultiIndex while keeping the information from it. Let's do a quick demo: ```python import pandas as pd cols = pd.MultiIndex.from_tuples([(0, 1), (0, 1)]) df = pd.DataFrame([[1,2], [3,4]], index=cols) ``` | | | 0 | 1 | | - | - | - | - | | 0 | 1 | 1 | 2 | | 1 | 3 | 4 | | After reset of the MultiIndex we will get: ```python df.reset_index() ``` | | level\_0 | level\_1 | 0 | 1 | | - | -------- | -------- | - | - | | 0 | 0 | 1 | 1 | 2 | | 1 | 0 | 1 | 3 | 4 | ### Reset single level of MultiIndex ```python df.reset_index(level=1) ``` ## Typical Errors on Pandas Drop MultiIndex When the droplevel is invoked on wrong axis: columns or rows like: ```python cols = pd.MultiIndex.from_tuples([(0, 1), (0, 1)]) df = pd.DataFrame([[1,2], [3,4]], index=cols) df.columns.droplevel() ``` error is raised: > ValueError: Cannot remove 1 levels from an index with 1 levels: at least one level must be left. Another example is when the levels are less the one which should be dropped: ```python cols = pd.MultiIndex.from_tuples([(0, 1), (0, 1)]) df = pd.DataFrame([[1,2], [3,4]], index=cols) df.columns.droplevel(level=3) ``` > IndexError: Too many levels: Index has only 1 level, not 4 ## Resources - [Pandas drop MultiIndex](https://github.com/softhints/datascientyst/blob/master/multiindex/2.pandas-drop-multiindex.ipynb?ref=datascientyst.com) - [pandas.MultiIndex.droplevel](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.MultiIndex.droplevel.html?ref=datascientyst.com) - [1pandas.DataFrame.reset\_index](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.reset%5Findex.html?ref=datascientyst.com) ### What is a DataFrame MultiIndex in Pandas URL: https://datascientyst.com/dataframe-multiindex-in-pandas/ Last updated: 2021-11-12T12:45:48.000Z **In Pandas MultiIndex is advanced indexing techniques for DataFrames**. It allows multiple levels for the indexes. It can be called also - **hierarchical index or multi-level index.** In Pandas indexes are represented as a labeled axis stored as an object. They help for: - identify data - data alignment - get and set of subsets ## Step 1: DataFrame MultiIndex in Pandas Let's create a MultiIndex DataFrame in Pandas. The first example shows MultiIndex for the rows: ```python cols = pd.MultiIndex.from_tuples([(0, 1), (0, 1)]) pd.DataFrame([[1,2], [3,4]], index=cols) ``` which will produce DataFrame with two level index: - 0 is the level 0 - 1 is the level 1 Note: In this case 0 and 1 are the same in the levels which are not mandatory. | | | 0 | 1 | | - | - | - | - | | 0 | 1 | 1 | 2 | | 1 | 3 | 4 | | The next example will create a hierarchical index for the columns. As you can see there are two levels of the column index: - a - which is level - 0 - b - is considered as level 1 ```python cols = pd.MultiIndex.from_tuples([('a', 'b'), ('a', 'b')]) pd.DataFrame([[1,2], [3,4]], columns=cols) ``` | | a | | | - | - | - | | | b | b | | 0 | 1 | 2 | | 1 | 3 | 4 | ## Step 2: Pandas MultiIndex to single index In this step will check how to convert a multi-level index to a single level one. This can be done by dropping levels from the MultiIndex: ```python df.droplevel(level=0, axis=1) ``` which will result in: | | b | b | | - | - | - | | 0 | 1 | 2 | | 1 | 3 | 4 | and ```python df.droplevel(level=1, axis=1) ``` which will result in: | | a | a | | - | - | - | | 0 | 1 | 2 | | 1 | 3 | 4 | Note 1: Rows or column can be selected by parameter - `axis` Note 2: If you try to use the labelled value: 'a' or 'b' like: `df.droplevel(level='b', axis=1)` you will get an error: > KeyError: 'Level b not found' ## Step 3: How to create Pandas MultiIndex There are many ways to create a DataFrame with Pandas MultiIndex. We already saw the one which uses the method: `from_tuples`. Now let's check another one - `from_arrays`: ```python mi = pd.MultiIndex.from_arrays( [[1, 2], [3, 4], [5, 6]], names=['x', 'y', 'z']) mi.to_frame() ``` which will result in: | | | | x | y | z | | - | - | - | - | - | - | | x | y | z | | | | | 1 | 3 | 5 | 1 | 3 | 5 | | 2 | 4 | 6 | 2 | 4 | 6 | ## Step 4: Convert Pandas MultiIndex into column Another useful operation for MultiIndex DataFrames is to convert levels into columns. This can be done in several ways: - `stack`/`unstack` - `pivot` - `crosstabs` - `reset_index` Let's check an example for: `reset_index`. So let's have the next DataFrame: ```python cols = pd.MultiIndex.from_tuples([('a', 'b'), ('a', 'b')]) df = pd.DataFrame([[1,2], [3,4]], index=cols) ``` | | | 0 | 1 | | - | - | - | - | | a | b | 1 | 2 | | b | 3 | 4 | | To convert MultiIndex to columns or rows can be done by: ```python df.reset_index() ``` which will output: | | level\_0 | level\_1 | 0 | 1 | | - | -------- | -------- | - | - | | 0 | a | b | 1 | 2 | | 1 | a | b | 3 | 4 | ![dataframe-multiindex-in-pandas](https://datascientyst.com/content/images/2021/07/dataframe-multiindex-in-pandas.png) ## Resources 1. [Pandas DataFrame MultiIndex Notebook](https://github.com/softhints/datascientyst/blob/master/multiindex/1.pandas-multiindex.ipynb?ref=datascientyst.com) 2. [pandas.MultiIndex](https://pandas.pydata.org/pandas-docs/stable/user%5Fguide/advanced.html?ref=datascientyst.com) ### How to Render Pandas DataFrame As HTML Table Keeping Style URL: https://datascientyst.com/render-pandas-dataframe-html-table-keeping-style/ Last updated: 2023-04-27T06:11:27.000Z In this guide, we'll show how to render Pandas DataFrame as a HTML table while keeping the style. We will cover striped tables and custom CSS formatting for Pandas DataFrames. If you like to find more advanced Pandas styling check: - [Pandas Visualization & Styling](https://datascientyst.com/visualization/) - [How to Set Pandas DataFrame Background Color Based On Condition/Value or Alternate Row Color based on Group ](https://datascientyst.com/pandas-dataframe-background-color-based-condition-value-alternate-row-color-based-group/) ## Step 1: Create DataFrame For this post we are going to use [DataFrame with random generated integers](https://datascientyst.com/how-to-create-a-dataframe-of-random-integers-with-pandas/). ```python df = pd.DataFrame({'a': {0: 3, 1: 3, 2: 4, 3: 0, 4: 0}, 'b': {0: 4, 1: 1, 2: 4, 3: 0, 4: 1}, 'c': {0: 0, 1: 2, 2: 3, 3: 0, 4: 0}, 'd': {0: 1, 1: 3, 2: 4, 3: 1, 4: 1}, 'e': {0: 2, 1: 0, 2: 3, 3: 1, 4: 0}, 'f': {0: 4, 1: 2, 2: 4, 3: 0, 4: 0}, 'g': {0: 3, 1: 1, 2: 2, 3: 4, 4: 0}}) ``` result: | | a | b | c | d | e | f | g | | - | - | - | - | - | - | - | - | | 0 | 3 | 4 | 0 | 1 | 2 | 4 | 3 | | 1 | 3 | 1 | 2 | 3 | 0 | 2 | 1 | | 2 | 4 | 4 | 3 | 4 | 3 | 4 | 2 | | 3 | 0 | 0 | 0 | 1 | 1 | 0 | 4 | | 4 | 0 | 1 | 0 | 1 | 0 | 0 | 0 | ## Step 2: Convert Dataframe to HTML table In order to Convert Dataframe to HTML table we are going to use method - `to_html`: ```python print(df.to_html()) ``` which will result in: ```html
0 1 2 3 4
``` This will generate pretty basic HTML table without any formatting. ## Step 3: Pandas DataFrame as striped table If you like to get **striped table from DataFrame**(similar to the Jupyterlab formatting with alternate row colors then you can use Pandas method `to_html` and set classes - `table table-striped`: ```python print(df.to_html(classes='table table-striped text-center', justify='center')) ``` The generated HTML will require Bootstrap in order to visualize the table properly. Bootstrap can be add to your site by CDN [Bootstrap Getting started](https://getbootstrap.com/docs/3.3/getting-started/?ref=datascientyst.com): In the section of your site add: `` In the section of your site at the bottom add: `` ## Step 4: Render DataFrame to HTML table keeping custom CSS style Finally let **apply custom styling and convert the DataFrame as HTML table**. So any style applied to Pandas DataFrame can be saved as HTML code. This can be done with method `render` as follows: ```python def format_color_groups(df): colors = ['gold', 'lightblue'] x = df.copy() factors = list(x['b'].unique()) i = 0 for factor in factors: style = f'background-color: {colors[i]}' x.loc[x['b'] == factor, :] = style i = not i return x df.sort_values(by='b').style.apply(format_color_groups, axis=None) ``` which will result in: ![pandas-to-html-style](https://datascientyst.com/content/images/2023/04/pandas-to-html-style.png) | | a | b | c | d | e | f | g | | - | - | - | - | - | - | - | - | | 0 | 3 | 4 | 0 | 1 | 2 | 4 | 3 | | 1 | 3 | 1 | 2 | 3 | 0 | 2 | 1 | | 2 | 4 | 4 | 3 | 4 | 3 | 4 | 2 | | 3 | 0 | 0 | 0 | 1 | 1 | 0 | 4 | | 4 | 0 | 1 | 0 | 1 | 0 | 0 | 0 | ### How to Create a Pandas DataFrame of Random Integers URL: https://datascientyst.com/how-to-create-a-dataframe-of-random-integers-with-pandas/ Last updated: 2021-11-27T08:29:42.000Z In this quick guide, we're going to create a Pandas DataFrame of random integers with arbitrary length. ```python import numpy as np import pandas as pd import string string.ascii_lowercase n = 5 m = 7 cols = string.ascii_lowercase[:m] df = pd.DataFrame(np.random.randint(0, n,size=(n , m)), columns=list(cols)) df ``` Explanation: - n - number of the rows - m - number of the columns. The columns will be named with latin letters in lowercase - `np.random.randint` \- will be used to produce random integers in a range of n to m. The produced DataFrame with random integer numbers is: | | a | b | c | d | e | f | g | | - | - | - | - | - | - | - | - | | 0 | 2 | 3 | 1 | 0 | 3 | 2 | 1 | | 1 | 1 | 4 | 3 | 3 | 3 | 0 | 4 | | 2 | 4 | 2 | 3 | 2 | 2 | 2 | 3 | | 3 | 3 | 1 | 1 | 3 | 2 | 2 | 4 | | 4 | 3 | 0 | 0 | 3 | 0 | 3 | 2 | ### Python Find All Parent/Child Nodes in Pandas DataFrame - Listing Subtree Descendants URL: https://datascientyst.com/python-find-all-parent-child-nodes-pandas-dataframe-subtree/ Last updated: 2021-06-22T10:07:50.000Z In this quick article, we'll have a look at **how to list all parent/child nodes in Pandas DataFrame.** This might be useful in certain scenarios like verifying trees or networks. ## Step 1: Prepare Hierarchical data Lets prepare hierarchical data which is going to be used for our example: ```python import pandas as pd df = pd.DataFrame({ 'parent': [0, 0, 1, 1, 2, 2, 3, 3, 4, 4], 'child': [1, 2, 3, 4, 5, 6, 6, 7, 8, 9] }) ``` Sometimes parent or child information might be stored in the index. In those cases index can be transfered to a column by: ```python df_req['index1'] = df_req.index ``` ## Step 2: Install Python package networkx As a second step we need to install package: `networkx` by: ```python pip install networkx ``` This package is used for creation, manipulation, and analysis of the structure and features of complex networks. ## Step 3: Listing Subtree Descendants with NetworkX in Pandas DataFrame Finally if we like to get all descendants of node 1 in this DataFrame we can do it by **converting the DataFrame records to NetworkX nodes**: ```python import networkx as nx g=nx.DiGraph() g.add_edges_from(df[['parent', 'child']].to_records(index=False)) ``` and then listing the subtree of NetworkX by: ```python from networkx.algorithms.traversal.depth_first_search import dfs_tree x = dfs_tree(g, 1) x.edges() ``` Which will result in: `OutEdgeView([(1, 3), (1, 4), (3, 6), (3, 7), (4, 8), (4, 9)])` Visually the same can be represented by: ![pandas-list-subtreee-dataframe](https://datascientyst.com/content/images/2021/06/pandas-list-subtreee-dataframe.png) ## Step 4: Set List of Descendants for each Row(Optional) In this step we are going to add a new column with a list of all descendants recursively. ```python def get_descendants(parent): descendants = list(dfs_tree(g, parent).edges()) return [x[1] for x in descendants] df["descendants"] = df["parent"].apply(get_descendants) ``` This will create new column with a list of all childs for the current parent: | | parent | child | descendants | | - | ------ | ----- | -------------------------------- | | 0 | 0 | 1 | \[1, 3, 6, 7, 4, 8, 9, 2, 5, 6\] | | 1 | 0 | 2 | \[1, 3, 6, 7, 4, 8, 9, 2, 5, 6\] | | 2 | 1 | 3 | \[3, 6, 7, 4, 8, 9\] | | 3 | 1 | 4 | \[3, 6, 7, 4, 8, 9\] | | 4 | 2 | 5 | \[5, 6\] | | 5 | 2 | 6 | \[5, 6\] | | 6 | 3 | 6 | \[6, 7\] | | 7 | 3 | 7 | \[6, 7\] | | 8 | 4 | 8 | \[8, 9\] | | 9 | 4 | 9 | \[8, 9\] | If you like to use custom code version you can use the one below. The code is a bit slower than the previous version: ```python def get_children(parent_id): list_of_children = [] def dfs(parent_id): child_ids = df[df["parent"]==parent_id]["child"] if child_ids.empty: return for child_id in child_ids: list_of_children.append(child_id) dfs(child_id) dfs(parent_id) return list_of_children df["descendants"] = df["parent"].apply(get_children) ``` ### How to Set Pandas DataFrame Background Color Based On Condition/Value or Alternate Row Color based on Group URL: https://datascientyst.com/pandas-dataframe-background-color-based-condition-value-alternate-row-color-based-group/ Last updated: 2022-10-28T06:05:52.000Z Using Pandas, we usually have many ways to group and sort values based on condition. In this short tutorial, we'll see how to set the background color of rows based on cell values from the cell row. In other words we are going to use a column on which to group and then apply styling as shown below: ![pandas-dataframe-background-color-based-condition-value-python](https://datascientyst.com/content/images/2022/10/pandas-dataframe-background-color-based-condition-value-python.webp) Whole code: ```python def format_color_groups(df): colors = ['gold', 'lightblue'] x = df.copy() factors = list(x['publication'].unique()) i = 0 for factor in factors: style = f'background-color: {colors[i]}' x.loc[x['publication'] == factor, :] = style i = not i return x df.style.apply(format_color_groups, axis=None) ``` ## Step 1: Read the data from Kaggle In this example we are going to use data from Kaggle. If you like to learn more please refer to: [How to Search and Download Kaggle Dataset to Pandas DataFrame ](https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/) ```python import pandas as pd df = pd.read_csv('./data/medium_data.csv.zip').sample(15).sort_values(by='publication') ``` So we are getting 15 random records sorted by column `'publication'` ## Step 2: Build styling method for alternate coloring In this step we are going to define a method which is going to change the background of the DataFrame based on value(s) from a given row. ```python def format_color_groups(df): colors = ['gold', 'lightblue'] x = df.copy() factors = list(x['publication'].unique()) i = 0 for factor in factors: style = f'background-color: {colors[i]}' x.loc[x['publication'] == factor, :] = style i = not i return x ``` We are going to use two colors: - 'gold' - 'lightblue' Then we select all unique values for the grouping column: ```python factors = list(x['publication'].unique()) ``` Finally we iterate over the rows of the DataFrame and alternate the color for each group: ```python for factor in factors: style = f'background-color: {colors[i]}' x.loc[x['publication'] == factor, :] = style i = not i ``` New DataFrame with styles is built and returned. ## Step 3: Apply Alternate Row Color based on Group Finally we apply the styling on the whole DataFrame: ```python df.style.apply(format_color_groups, axis=None) ``` We will achive the next result: ![pandas-dataframe-background-color-based-condition-value-python](https://datascientyst.com/content/images/2021/06/pandas-dataframe-background-color-based-condition-value-python.png) ### How to Search and Download Kaggle Dataset to Pandas DataFrame URL: https://datascientyst.com/search-download-kaggle-dataset-pandas-dataframe/ Last updated: 2021-06-04T21:25:04.000Z In this post, we'll take a brief look at the [Kaggle Datasets](https://www.kaggle.com/datasets?ref=datascientyst.com) and how to download/import them with Python. By the end, we'll see how to list, download single or multiple datasets and finally how to read them into Pandas DataFrame. ## Step 1: Create Kaggle API token First you will need to visit: [Kaggle](https://www.kaggle.com/?ref=datascientyst.com) and create a new account. You can sign up with your google account. In order to create new Kaggle API token follow: - Open your profile picture(top right) - Account - the url is: `https://www.kaggle.com//account` - API - Create new API Token - This will generate `kaggle.json` file - Place the file in your home folder as: `~/.kaggle/kaggle.json` - For more security (optional) - `chmod 600 ~/.kaggle/kaggle.json` More info is available on this link: [Kaggle API](https://github.com/Kaggle/kaggle-api?ref=datascientyst.com) ## Step 2: Install Python's package for Kaggle Next we are going to install the package which is going to download the datasets from Kaggle. You can install kaggle package in virtual environment by: ```bash pip install kaggle ``` or for the user: ```bash pip install --user kaggle ``` ## Step 3: Download single file from Kaggle dataset Now we are going to demonstrate how to download a single CSV file from the Kaggle dataset. This will work only if previous steps were done successfully: ```python import kaggle kaggle.api.authenticate() kaggle.api.dataset_download_file('dorianlazar/medium-articles-dataset', file_name='medium_data.csv', path='data/') ``` In the example above we are going to download file: `medium_data.csv` from: [dorianlazar/medium-articles-dataset](https://www.kaggle.com/dorianlazar/medium-articles-dataset?select=medium%5Fdata.csv&ref=datascientyst.com). The file will be downloaded in the folder `data/`. The file can be read by: ```python import pandas as pd pd.read_csv('data/medium_data.csv.zip') ``` which produce: ![search-and-download-kaggle-dataset-to-pandas-dataframe](https://datascientyst.com/content/images/2021/06/search-and-download-kaggle-dataset-to-pandas-dataframe.png) ## Step 4: Download multiple files from Kaggle dataset If we like to get all files from a Kaggle dataset then we can get them by: ```python import kaggle kaggle.api.authenticate() kaggle.api.dataset_download_files('dorianlazar/medium-articles-dataset', path='data/') ``` Note that this might be pretty slow for big datasets. We are downloading all files from the dataset mentioned above. ## Step 5: List and search Kaggle datasets with API Finally let's find how to list and search for Kaggle datasets. This can be done by next command: ```bash !kaggle datasets list -s article ``` Where we are searching for the keyword - `article`. The output is: | ref | title | size | lastUpdated | downloadCount | voteCount | usabilityRating | | | ------------------------------------------------------------ | ------------------------------------------------ | ------ | -------------------- | -------------- | ---------- | ---------------- | | | \----------------------------------------------------------- | \----------------------------------------------- | \----- | \------------------- | \------------- | \--------- | \--------------- | | | dorianlazar/medium-articles-dataset | Medium articles dataset | 1GB | 2020-06-30 14:13:56 | 1804 | 103 | 0.9411765 | | | hsankesara/medium-articles | Medium Articles | 1MB | 2018-06-17 08:45:49 | 1983 | 72 | 0.88235295 | | | gspmoreira/articles-sharing-reading-from-cit-deskdrop | Articles sharing and reading from CI&T DeskDrop | 8MB | 2017-08-27 21:33:01 | 10062 | 135 | 0.8235294 | | | jkkphys/english-wikipedia-articles-20170820-sqlite | English Wikipedia Articles 2017-08-20 SQLite | 7GB | 2018-11-27 21:54:22 | 1417 | 84 | 0.875 | | | asad1m9a9h6mood/news-articles | News Articles | 2MB | 2017-04-30 11:02:29 | 2731 | 31 | 0.8235294 | | | residentmario/wikipedia-article-titles | Wikipedia Article Titles | 73MB | 2017-09-22 16:42:20 | 726 | 26 | 0.75 | | | abhishek/10k-german-news-articles | 10k German News Articles | 123MB | 2019-11-07 08:50:32 | 552 | 89 | 0.8235294 | | | yufengdev/bbc-fulltext-and-category | BBC articles fulltext and category | 2MB | 2018-06-08 05:44:22 | 3799 | 35 | 0.64705884 | | | danofer/dbpedia-classes | DBPedia Classes | 166MB | 2019-07-04 11:30:52 | 979 | 26 | 1.0 | | | vetrirah/janatahack-independence-day-2020-ml-hackathon | NLP on Research Articles | 11MB | 2020-08-19 14:35:13 | 309 | 24 | 1.0 | | | jkkphys/english-wikipedia-articles-20170820-models | English Wikipedia Articles 2017-08-20 Models | 925MB | 2018-11-28 17:09:32 | 379 | 19 | 0.8125 | | | blessondensil294/topic-modeling-for-research-articles | Topic Modeling for Research Articles | 11MB | 2020-08-18 08:53:26 | 321 | 21 | 1.0 | | | szymonjanowski/internet-articles-data-with-users-engagement | Internet news data with readers engagement | 3MB | 2020-11-21 17:09:57 | 4069 | 330 | 0.9411765 | | | maxscheijen/dutch-news-articles | Dutch News Articles | 135MB | 2021-05-24 08:01:12 | 104 | 13 | 1.0 | | | aiswaryaramachandran/medium-articles-with-content | Medium Articles (with Content) | 218MB | 2018-11-10 18:17:46 | 569 | 29 | 0.7352941 | | | jeet2016/us-financial-news-articles | US Financial News Articles | 1GB | 2018-09-05 01:27:43 | 2530 | 54 | 0.625 | | | urbanbricks/wikipedia-promotional-articles | Wikipedia Promotional Articles | 201MB | 2019-10-27 16:31:06 | 276 | 15 | 1.0 | | | hkapoor/indian-financial-news-articles-20032020 | Indian financial news articles (2003-2020) | 3MB | 2020-05-26 20:41:29 | 233 | 25 | 1.0 | | | naharrison/particle-identification-from-detector-responses | Particle Identification from Detector Responses | 83MB | 2018-10-24 21:14:33 | 298 | 28 | 0.7058824 | | | zshujon/40k-bangla-newspaper-article | 40k Bangla Newspaper Article | 64MB | 2018-09-22 09:54:40 | 276 | 10 | 0.5625 | | ### How to Plot Time Series As work timetable in Pandas URL: https://datascientyst.com/plot-time-series-work-timetable-pandas/ Last updated: 2021-06-03T21:52:50.000Z In this quick guide, we'll show how to **plot time series as work timetable in Pandas**. The process includes several steps in order to prepare the data. First let's start with generating time series data for the work timetable. Generating range of time series between 2 dates with frequency of 30 minutes: ```python periods = pd.date_range(start='1/1/2021', end='1/5/2021', freq='30min') df = pd.DataFrame({'start_time':periods[:-1]}) ``` Note: the `[:-1]` is excluding the single period generated for the last date. Next we are going to generate column - `availability` \- which will be two options: - 0: 'UNAVAILABLE' - 1:'AVAILABLE' ```python import numpy as np df['availability'] = np.random.randint(0, 2, df.shape[0]) df['availability'] = df['availability'].map({0: 'UNAVAILABLE', 1:'AVAILABLE'}) ``` The picture below shows the initial data and the final output: ![pandas-plot-time-series-work-timetable](https://datascientyst.com/content/images/2021/06/plot-time-series-work-timetable-pandas.png) ## Step 1: Split Datetime column into date and time columns To start we need to prepare our datetime column. First we will ensure that it's really a datetime column. Next we will split the `start_time` into two parts: - date - time This will help us to plot the timetable: ```python df['start_time'] = pd.to_datetime(df['start_time']) df['date'] = df['start_time'].dt.date df['start_time_only'] = df['start_time'].dt.time ``` ## Step 2: Map string values to numbers Depending on your data you may need to remap the values. If so, you can use a dictionary where the key is the old value. The new one is the value in the dictionary: ```python df['availability'] = df['availability'].map({'UNAVAILABLE': 0, 'AVAILABLE': 1}) ``` ## Step 3: Convert time series to Work Timetable Finally we are going to generate the timetable itself by: ```python df[['availability', 'date','start_time_only']].set_index(['date','start_time_only']).unstack().swaplevel(0, 1, axis = 1).sort_index(axis = 1).T.style.background_gradient(cmap='ocean_r') ``` Let's add some explanation to the code above: - we select only few columns from the DataFrame (in case of many) - set new index by `.set_index(['date','start_time_only'])` \- so basically we leave only 1 column - `availability` - `.unstack()` will pivot the hierarchical index labels - `.swaplevel(0, 1, axis = 1).sort_index(axis = 1).T` \- is optional. This will swap the levels, sort the index(if needed) and finally will transpose the rows to columns - `.style.background_gradient(cmap='ocean_r')` \- finally we are going to add style in order to highlight the options On the picture below you can follow the steps: ![pandas-plot-time-series-work-timetable-all-steps](https://datascientyst.com/content/images/2021/06/pandas-plot-time-series-work-timetable-all-steps.png) Below you can find the final timetable: ![pandas-plot-time-series-work-timetable](https://datascientyst.com/content/images/2021/06/pandas-plot-time-series-work-timetable.png) ### How to Fix Pandas to_datetime: Wrong Date and Errors URL: https://datascientyst.com/how-to-fix-pandas-to_datetime-wrong-date-and-errors/ Last updated: 2022-08-29T07:52:50.000Z In this article, we'll learn how to **convert dates saved as strings in Pandas with method [to\_datetime](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.to%5Fdatetime.html?ref=datascientyst.com)**. We will cover the basic usage, problematic dates and errors. For this article we are going to generate dates with the code below: ## 1: Convert Strings to Datetime in Pandas DataFrame First let's show how to **convert a list of dates stored as strings to datetime in a DataFrame**. ```python import pandas as pd df = pd.DataFrame({'date_str':['05/18/2021', '05/19/2021', '05/20/2021']}) ``` We have a DataFrame with 1 column with several dates. The conversion to datetime column is done by: ```python df['date'] = pd.to_datetime(df['date_str']) ``` We need to provide the column and the method will convert anything which is like a date. **Note:** To check the column type, is it a string or datetime, we are going to use: `df.dtypes`. This results into: ``` date_str object date datetime64[ns] ``` ## 2: Typical Errors with Pandas to\_datetime In case of errors you will get: `ParserError`. Simple example: ```python df = pd.DataFrame({'date_str':['0']}) ``` This will bring error: > ParserError: day is out of range for month: 0 or another frequent error is produced by: ```python df = pd.DataFrame({'date_str':['a']}) ``` results in: > ValueError: Given date string not likely a datetime. Those errors are resolved by adding parameter - `errors`: ```python pd.to_datetime(df['date_str'], errors='coerce') ``` or ```python pd.to_datetime(df['date_str'], errors='ignore') ``` Where options are: - errors : {'ignore', 'raise', 'coerce'}, default 'raise' - If 'raise', then invalid parsing will raise an exception. - If 'coerce', then invalid parsing will be set as NaT. - If 'ignore', then invalid parsing will return the input. ## 3: Fix Pandas to\_datetime produces wrong dates Sometimes you will end with successful conversion without Python errors. Yet you may face unexpected results. ### Single format wrong dates With the code below we are going to generate 50 dates with the same format: `%d/%m/%Y`: ```python def make_dates(): dates = [] for i in range(50, 0, -1): day = (datetime.now() - timedelta(days=i)).date().strftime('%d/%m/%Y') dates.append(day) return dates ``` Basic conversion of the those dates will produce unexpected results: ```python df['date'] = pd.to_datetime(df['date_str']) df.groupby(['date']).date_str.count().plot(kind='bar', figsize=(20,5)) ``` ![pandas-to_datetime-wrong-date-conversion](https://datascientyst.com/content/images/2021/06/pandas-to_datetime-wrong-date-conversion.png) So `07/05/2021` is treated as `05/07/2021`. The same is for `08/05/2021`, `09/05/2021`. To fix this date parsing problems we can use next syntax: `pd.to_datetime(df['date_str'], format='%d/%m/%Y')` which will produce correct datetime conversion by forcing the date format: ![pandas-to_datetime-wrong-date-format](https://datascientyst.com/content/images/2021/06/pandas-to_datetime-wrong-date-format.png) ### Two formats wrong dates We are going to produce a list of dates for the last 30 days. First 15 will be with format: `'%m/%d/%Y'` and second half will be `'%d/%m/%Y'`: ```python from datetime import datetime, timedelta def make_dates(): dates = [] for i in range(15, 0, -1): day = (datetime.now() - timedelta(days=i)).date().strftime('%m/%d/%Y') dates.append(day) for i in range(30, 15, -1): day = (datetime.now() - timedelta(days=i)).date().strftime('%d/%m/%Y') dates.append(day) return dates ``` The expected output of the code is: \['05/18/2021', '05/19/2021', '05/20/2021', '05/21/2021', ... '15/05/2021', '16/05/2021', '17/05/2021'\] If we try direct conversion of the DataFrame above we might get unexpected results showing the plot below: ![pandas-to_datetime-wrong-date-and-errors](https://datascientyst.com/content/images/2021/06/pandas-to_datetime-wrong-date-and-errors.png) The problem above will be the result of mixed formats which is not obvious. You may need to analyze your data before conversion. In next section is the solution of this problem ## 4: Pandas to\_datetime convert several formats Let's continue from the last section and convert the same DataFrame with two and more date formats. Below you can see how to convert string dates when the formats are clear. First we are going to parse the first format with `format='%d/%m/%Y', errors='coerce'` and then the second one will be processed plus mask: ```python df['date'] = pd.to_datetime(df['date_str'], format='%d/%m/%Y', errors='coerce') mask = df['date'].isnull() df.loc[mask, 'date'] = pd.to_datetime(df['date_str'], format='%m/%d/%Y', errors='coerce') ``` The dates from the both formats will be correctly parsed: ![pandas-to_datetime-multi-format-date-parsing](https://datascientyst.com/content/images/2021/06/pandas-to_datetime-multi-format-date-parsing.png) ## Conclusion In this short tutorial, we looked at a general use case of using *to\_datetime*. We also looked at the reasons for typical errors using *to\_datetime* with bad data. Finally we covered how to analyse datetime columns and how to convert mixed date formats. ### How to reshape Pandas Series into 2d array URL: https://datascientyst.com/reshape-pandas-series-into-2d-array/ Last updated: 2022-10-30T06:17:10.000Z In this article, we'll talk about **how to reshape Pandas Series into 2d array**. The example is inspired by getting data from a web/HTML table. Copying a web table might come as a single column instead of a table. ![reshape-pandas-series-into-2d-array](https://datascientyst.com/content/images/2022/10/reshape-pandas-series-into-2d-array.webp) Suppose you need to copy paste next table: | **Distribution of points in ICC World Test Championship** | | | | | | --------------------------------------------------------- | -------------------- | -------------------- | --------------------- | ----------------------- | | **Matches in series** | **Points for a win** | **Points for a tie** | **Points for a draw** | **Points for a defeat** | | 2 | 60 | 30 | 20 | 0 | | 3 | 40 | 20 | 13 | 0 | | 4 | 30 | 15 | 10 | 0 | | 5 | 24 | 12 | 8 | 0 | Instead of getting a nice table you will end with a single column in form of: ``` Matches in series Points for a win Points for a tie ``` ## Step 1: Get the data as a single column DataFrame ### Method 1 The first step is to copy paste the table data into a new CSV file and save it as csv - `data.csv`. Then read it by: ```python import pandas as pd df = pd.read_csv('~/Desktop/data.csv') ``` The second way is by reading the table from web address: ```python import pandas as pd df = pd.read_html('https://blog.softhints.com/how-to-merge-multiple-csv-files-with-python/') ``` The final DataFrame should look like: | | Distribution of points in ICC World Test Championship | | - | ----------------------------------------------------- | | 0 | Matches in series | | 1 | Points for a win | | 2 | Points for a tie | | 3 | Points for a draw | | 4 | Points for a defeat | ## Step 2: Find the shape for the initial table In this step we need to find what is the shape which should be used for reshaping. In case of incorrect shape error is raised: > ValueError: cannot reshape array of size 25 into shape (4) In order to do this we need to do two steps: - Check the original table - it has 5 columns - then we can get the names - Check the DataFrame data by: ```python df.shape ``` To get if there is (25, 1) The number is divisible to 5 which is the expected. If this is not the case we have extra data - empty lines, formatting which needs to be cleaned. ## Step 3: Reshape Series - convert single column to multiple columns To reshape Series to a DataFrame which has the same table form as the original source: ```python pd.DataFrame(df.iloc[5:, :].values.reshape(-1, 5), columns=df.iloc[:5, 0].values) ``` Which give us result of: | | Matches in series | Points for a win | Points for a tie | Points for a draw | Points for a defeat | | - | ----------------- | ---------------- | ---------------- | ----------------- | ------------------- | | 0 | 2 | 60 | 30 | 20 | 0 | | 1 | 3 | 40 | 20 | 13 | 0 | | 2 | 4 | 30 | 15 | 10 | 0 | | 3 | 5 | 24 | 12 | 8 | 0 | Let's explain how reshaping is working: - `df.iloc[5:, :]` \- we read all values except the first 5\. Since they are the table header we don't need them. - `.reshape(-1, 5)` \- next we change the shape of the values ( from previous piece of code (20, 1) to (-1, 5). This means making 5 columns and the required number of rows. You can explicitly write number of rows as - (4, 5) - `columns=df.iloc[:5, 0].values` \- as you already guess - get the header values and set them as column names