Skip to content

Gaussian kde contour univariate

daspi.plotlib.plotter.GaussianKDEContourUnivariate(source, target, feature, width=CATEGORY.FEATURE_SPACE, skip_na=None, fill=True, fade_outers=True, n_points=DEFAULT.KD_SEQUENCE_LEN, target_on_y=True, color=None, ax=None, visible_spines=None, hide_axis=None, **kwds)

Bases: TransformPlotter

Class for creating univariate contour plotters. This is a special case of the GaussianKDEContour plot, where the contour lines are plotted on top of each other, resulting in a univariate plot with contour lines.

This plot can be used to show the distribution of a univariate data set in a more detailed way than the GaussianKDE plot. The contour lines represent different levels of density, which can be highlighted using different colors or opacities.

PARAMETER DESCRIPTION
source

Pandas long format DataFrame containing the data source for the plot.

TYPE: pandas DataFrame

target

Column name of the target variable for the plot.

TYPE: str

feature

Column name of the feature variable for the plot, by default ''

TYPE: str

fill

Flag indicating whether to fill between the contour lines, by default True

TYPE: bool DEFAULT: True

fade_outers

Flag indicating whether the outer lines of the contour plot should be faded. This has no effect if fill is True, by default True.

TYPE: bool DEFAULT: True

n_points

Number of points the estimate and the sequence should have. Note that the calculated points are equal to the square of the given number (because the contour is two-dimensional). by default KD_SEQUENCE_LEN (defined in constants.py)

TYPE: int DEFAULT: KD_SEQUENCE_LEN

margin

Margin for the sequence as factor of data range, by default 0.2.

TYPE: float

target_on_y

Flag indicating whether the target variable is plotted on the y-axis. If False, all contour lines have the same color. by default True

TYPE: bool DEFAULT: True

color

Color to be used to draw the artists. If None, the first color is taken from the color cycle, by default None.

TYPE: str | None DEFAULT: None

ax

The axes object for the plot. If None, the current axes is fetched using plt.gca(). If no axes are available, a new one is created. Defaults to None.

TYPE: Axes | None DEFAULT: None

visible_spines

Specifies which spines are visible, the others are hidden. If 'none', no spines are visible. If None, the spines are drawn according to the stylesheet. Defaults to None.

TYPE: Literal['target', 'feature', 'none'] | None DEFAULT: None

hide_axis

Specifies which axes should be hidden. If None, both axes are displayed. Defaults to None.

TYPE: Literal['target', 'feature', 'both'] | None DEFAULT: None

**kwds

Those arguments have no effect. Only serves to catch further arguments that have no use here (occurs when this class is used within chart objects).

DEFAULT: {}

Examples:

Apply to an existing Axes object:

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from daspi import GaussianKDEContourUnivariate

fig, ax = plt.subplots()
df = pd.DataFrame(dict(
    category = ['A'] * 50 + ['B'] * 50 + ['C'] * 50,
    value = (
        list(np.random.normal(loc=10, scale=2, size=50))
        + list(np.random.normal(loc=15, scale=2, size=50))
        + list(np.random.normal(loc=12, scale=2, size=50)))))
plotter = GaussianKDEContourUnivariate(
    source=df, target='value', feature='category', 
    fill=True, n_points=50, ax=ax)
plotter()
plotter.label_feature_ticks()

Apply using the plot method of a DaSPi Chart object:

import daspi as dsp

df = dsp.load_dataset('painkillers-dissolution')
chart = dsp.SingleChart(
        source=df,
        target='time',
        feature='brand',
        categorical_feature=True,
    ).plot(
        dsp.GaussianKDEContourUnivariate,
        fill=True,
        n_points=50
    ).label(
        feature_label='Brand',
        target_label='Time (s)'
    )

With hue grouping for multiple colors:

import daspi as dsp

df = dsp.load_dataset('painkillers-dissolution')
chart = dsp.SingleChart(
        source=df,
        target='time',
        feature='brand',
        hue='stirrer',
        dodge=True,
    ).plot(
        dsp.GaussianKDEContourUnivariate,
        fill=True,
        fade_outers=True,
        n_points=50
    ).label(
        feature_label='Brand',
        target_label='Time (s)'
    )

Comparison with Violin plot in a JointChart:

import daspi as dsp

df = dsp.load_dataset('painkillers-dissolution')
chart = dsp.JointChart(
        source=df,
        target='time',
        feature='brand',
        hue='stirrer',
        ncols=1,
        nrows=2,
        sharex=True,
        dodge=(False, True),
        target_on_y=True
    ).plot(
        dsp.GaussianKDEContourUnivariate,
        fill=True,
        n_points=50
    ).plot(
        dsp.Violin
    ).label(
        feature_label=(True, True),
        target_label=(True, True)
    )

n_points = n_points instance-attribute

Number of points the estimate and the sequence should have.

shape = (n_points, n_points) instance-attribute

Shape used to reshape data before plotting the contours.

fill = fill instance-attribute

Flag indicating whether to fill between the contour lines.

width = width instance-attribute

The maximum width of the contour.

cmap = LinearSegmentedColormap.from_list('', colors) instance-attribute

The colormap to be used for the contour plot.

kw_default property

Return the default keyword arguments for the plot.

transform(feature_data, target_data)

Perform the transformation on the target data by estimating its 2D kernel density. Feature data is generated with a gaussian distribution centered at feature_data with width as std, and target data is sorted to align with this distribution.

PARAMETER DESCRIPTION
feature_data

Base location (offset) of feature axis coming from feature_grouped generator.

TYPE: float | int

target_data

Target data for which the kernel density is estimated.

TYPE: pandas Series

RETURNS DESCRIPTION
pandas DataFrame

Transformed data with estimated 2D kernel density.