# Ignore "Non-leaves rows" for sunburst diagram?

**URL:** <https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789>\
**Category:** 📊 Plotly Python\
**Tags:** question\
**Created:** [February 7, 2022, 3:43pm UTC](https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789 "2022-02-07T15:43:09Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![edent](https://sea2.discourse-cdn.com/flex024/user_avatar/community.plotly.com/edent/32/17938_2.png) [@edent](https://community.plotly.com/u/edent)\
**Post date:** [February 7, 2022, 3:43pm UTC](https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789/1 "2022-02-07T15:43:09Z")

</div>

I have a DataFrame of hierarchical data:

```auto
       0 1 2 3 4
0 alice bob chuck david ella
0 alice bob chuck david fred
0 alice bob chuck NaN NaN

```

If I try to create a Sunburst plot, I get told

> Non-leaves rows are not permitted in the dataframe

The same thing occurs if I replace the `NaN` with `None`

I’m aware that I could replace the `NaN`s with some dummy text, but that will distort the diagram I’m drawing.

Is there a way to skip these non-leave rows?

Thanks!

---

<div class="post-metadata">

**Author:** ![edent](https://sea2.discourse-cdn.com/flex024/user_avatar/community.plotly.com/edent/32/17938_2.png) [@edent](https://community.plotly.com/u/edent)\
**Post date:** [February 7, 2022, 4:17pm UTC](https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789/2 "2022-02-07T16:17:17Z")

</div>

I commented out the check in `plotly/express/_core.py` and it worked.

See:

```auto
def _check_dataframe_all_leaves(df):
    df_sorted = df.sort_values(by=list(df.columns))
    null_mask = df_sorted.isnull()
    df_sorted = df_sorted.astype(str)
    null_indices = np.nonzero(null_mask.any(axis=1).values)[0]
    for null_row_index in null_indices:
        row = null_mask.iloc[null_row_index]
        i = np.nonzero(row.values)[0][0]
        if not row[i:].all():
            raise ValueError(
                "None entries cannot have not-None children",
                df_sorted.iloc[null_row_index],
            )
    df_sorted[null_mask] = ""
    row_strings = list(df_sorted.apply(lambda x: "".join(x), axis=1))
    #for i, row in enumerate(row_strings[:-1]):
        #if row_strings[i + 1] in row and (i + 1) in null_indices:
            #raise ValueError(
            # "Non-leaves rows are not permitted in the dataframe \n",
            # df_sorted.iloc[i + 1],
            # "is not a leaf.",
            #)

```

My diagram renders perfectly. It would be great if there was an `ignore_non_leaves=True` option, rather than my horrible hack!

---

<div class="post-metadata">

**Author:** ![ddavo](https://sea2.discourse-cdn.com/flex024/user_avatar/community.plotly.com/ddavo/32/18123_2.png) [@ddavo](https://community.plotly.com/u/ddavo)\
**Post date:** [February 28, 2022, 10:59am UTC](https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789/3 "2022-02-28T10:59:21Z")

</div>

These non-leave rows are automatically created by Plotly, so you can safely delete them

Instead of commenting code from the library (which won’t work if you update or go to another machine), you can modify the dataframe:

```auto
df = df.dropna()

```

---

<div class="post-metadata">

**Author:** ![yhs](https://sea2.discourse-cdn.com/flex024/user_avatar/community.plotly.com/yhs/32/18820_2.png) [@yhs](https://community.plotly.com/u/yhs)\
**Post date:** [April 13, 2022, 7:13am UTC](https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789/4 "2022-04-13T07:13:49Z")

</div>

Thanks @ddavo , you are the man!

---

<div class="post-metadata">

**Author:** ![ameyakambli](https://avatars.discourse-cdn.com/v4/letter/a/a9a28c/32.png) [@ameyakambli](https://community.plotly.com/u/ameyakambli)\
**Post date:** [April 21, 2022, 1:26am UTC](https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789/5 "2022-04-21T01:26:22Z")

</div>

> [@ddavo](#):
>
> ated by Plotly, so you can safely delete them
> 
> Instead of

But won’t this cause an issue as both column 3 and 4 would dropped if we use df.dropna()

---

<div class="post-metadata">

**Author:** ![ddavo](https://sea2.discourse-cdn.com/flex024/user_avatar/community.plotly.com/ddavo/32/18123_2.png) [@ddavo](https://community.plotly.com/u/ddavo)\
**Post date:** [April 21, 2022, 8:21am UTC](https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789/6 "2022-04-21T08:21:40Z")

</div>

dropna() by default will drop rows with Null values, not columns

---

<div class="post-metadata">

**Author:** ![ameyakambli](https://avatars.discourse-cdn.com/v4/letter/a/a9a28c/32.png) [@ameyakambli](https://community.plotly.com/u/ameyakambli)\
**Post date:** [April 22, 2022, 7:54pm UTC](https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789/8 "2022-04-22T19:54:16Z")

</div>

```auto
> 0 1 2 3 4
0 alice bob smith NaN NaN
0 alice bob smith rocky NaN
0 alice bob chuck david ella
0 alice bob chuck david fred
0 alice bob chuck NaN NaN

```

This is more specific to my current scenario, by doing dropna() i will lose rows branching out from bob to smith. Smith and Chuck are two children belonging to Bob, how do i visualize a treemap/sunburst in such scenario without filling NaN with some dummy text

---

<div class="post-metadata">

**Author:** ![ddavo](https://sea2.discourse-cdn.com/flex024/user_avatar/community.plotly.com/ddavo/32/18123_2.png) [@ddavo](https://community.plotly.com/u/ddavo)\
**Post date:** [April 23, 2022, 9:10am UTC](https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789/9 "2022-04-23T09:10:08Z")

</div>

This is because the DataFrame expected needs to be _rectangular_. With each column having a value to group by. In the [docs](https://plotly.com/python/treemaps/), you have an example with a dataframe grouped by day, then by time and then by sex.

![image](https://us1.discourse-cdn.com/flex024/uploads/plot/original/3X/0/f/0fe6a93f6557922eeb31ebef926c9aedc3ab5cce.png)

In your case, I think it would be easier to just use `names` and `parents` instead of passing a DataFrame.

```auto
fig = px.treemap(
    names = ["Alice", "Bob" , "Smith", "Rocky", "Chuck", "David", "Ella" , "Fred"],
    parents = ["" , "Alice", "Bob" , "Smith", "Bob" , "Chuck", "David", "David"]
)

```

---

<div class="post-metadata">

**Author:** ![noLogoInTheFoam](https://avatars.discourse-cdn.com/v4/letter/n/e36b37/32.png) [@noLogoInTheFoam](https://community.plotly.com/u/noLogoInTheFoam)\
**Post date:** [December 1, 2022, 5:39pm UTC](https://community.plotly.com/t/ignore-non-leaves-rows-for-sunburst-diagram/60789/11 "2022-12-01T17:39:35Z")

</div>

Thanks @edent! I’ve also had to use this hack.
