Q-Logo 我的学习笔记分享

Pandas DataFrame的round函数详解及小坑

官方文档

Pandas DataFrame提供了round函数,可以将DataFrame中各列舍入到不同精度,其官方文档用法说明如下:

df.round(decimals=0, *args, **kwargs)

输入参数

decimals : int, dict, Series,每一列舍入到的小数位数。如果是整数,每一列都被舍入到这个位数;如果是字典或序列,各列舍入到指定的精度。列的名字应该作为decimals 字典的键,或者decimals 序列的index。未在decimals 中指定精度的列将保留原样。如果decimals 中有不是列名的键或index,会被忽略。

返回:

DataFrame: 舍入到指定精度的DataFrame。

举例:

>>> df = pd.DataFrame(np.random.random([3, 3]),
... columns=['A', 'B', 'C'], index=['first', 'second', 'third'])
>>> df
A B C
first 0.028208 0.992815 0.173891
second 0.038683 0.645646 0.577595
third 0.877076 0.149370 0.491027
>>> df.round(2)
A B C
first 0.03 0.99 0.17
second 0.04 0.65 0.58
third 0.88 0.15 0.49
>>> df.round({'A': 1, 'C': 2})
A B C
first 0.0 0.992815 0.17
second 0.0 0.645646 0.58
third 0.9 0.149370 0.49
>>> decimals = pd.Series([1, 0, 2], index=['A', 'B', 'C'])
>>> df.round(decimals)
A B C
first 0.0
1 0.17
second 0.0
1 0.58
third 0.9
0 0.49

官方文档对round函数的说明很清楚,举例也很全面。

但是,这里仍然有一些小坑,这里扒出来看一看。

decimals=0时,返回的是浮点数而非整数

上面举的最后一个例子中,给出的返回值,B 列的值是1, 1, 0,但如果照着例子去实际操作下来,得到的结果却是 1.0, 1.0, 0.0 。即最后一个例子中df.round(decimals)显示的实际结果是:

>>> df.round(decimals)
A B C
first 0.0
1.0 0.17
second 0.0
1.0 0.58
third 0.9
0.0 0.49

这表明decimals = 0的情况下,round 的结果是浮点数,而非整数。

如果就是想得到整数,可以这样:

>>> df2 = df.round(decimals)

>>> df2['B'] = df2['B'].astype(int)

>>> df2

A B C
first 0.0 1 0.17
second 0.0 1 0.58
third 0.9 0 0.49

round 并非四舍五入,对浮点数执行round 须谨慎

Pandas DataFrame中的round 函数,实际上调用了numpy 中的around 函数, 此函数的文档中提到

For values exactly halfway between rounded decimal values, NumPy
rounds to the nearest even value. Thus 1.5 and 2.5 round to 2.0,
-0.5 and 0.5 round to 0.0, etc. Results may also be surprising due
to the inexact representation of decimal fractions in the IEEE
floating point standard [1]_ and errors introduced when scaling
by powers of ten.

也就是说,其舍入的策略是“四舍六入五成双”,但由于计算机存储的是二进制数,在将输入的十进制浮点小数转成二进制小数保存时,可能并不精确,因此某些时候会产生不符合我们预期的结果。这与python 的内置round 函数是一样的,想要详细了解可以参考我的另一篇文章从Python 的round 函数的坑谈四舍五入。