|
quickstats
Quickly compute simple statistics
|
Quickly compute simple statistics. More...
Classes | |
| struct | MadOptions |
Options for mad() and mad_with_infinities(). More... | |
| struct | MedianOptions |
Options for median(). More... | |
| class | MultipleQuantilesFixedNumber |
| Calculate multiple quantiles from a fixed number of elements. More... | |
| class | MultipleQuantilesVariableNumber |
| Calculate multiple quantiles from a variable number of elements. More... | |
| struct | MultipleQuantilesVariableNumberOptions |
Options for MultipleQuantilesVariableNumber. More... | |
| struct | PairwiseSumOptions |
Options for pairwise_sum() and pairwise_sum_abstract(). More... | |
| struct | PairwiseSumWorkspace |
Re-usable workspace for pairwise_sum() and pairwise_sum_abstract(). More... | |
| struct | RssOptions |
Options for rss(). More... | |
| struct | RssResult |
Result of rss(). More... | |
| struct | RssWorkspace |
Re-usable workspace for rss(). More... | |
| class | SingleQuantileFixedNumber |
| Calculate a single quantile from a fixed number of elements. More... | |
| class | SingleQuantileVariableNumber |
| Calculate a quantile from a variable number of elements. More... | |
| struct | SingleQuantileVariableNumberOptions |
Options for SingleQuantileVariableNumber. More... | |
Functions | |
| template<typename Output_ = double, typename Input_ > | |
| Output_ | mad (const std::size_t num_total, Input_ *const ptr, const Input_ median, const MadOptions< Output_ > &options) |
| template<typename Output_ = double, typename Input_ > | |
| Output_ | mad (const std::size_t num_total, const std::size_t num_non_zero, Input_ *const values, const Input_ median, const MadOptions< Output_ > &options) |
| template<typename Float_ > | |
| Float_ | scale_mad_to_sd (const Float_ x) |
| template<typename Output_ = double, typename Input_ > | |
| Output_ | median (const std::size_t num_total, Input_ *const ptr, const MedianOptions< Output_ > &options) |
| template<typename Output_ = double, typename Input_ > | |
| Output_ | median (const std::size_t num_total, const std::size_t num_non_zero, Input_ *const values, const MedianOptions< Output_ > &options) |
| template<std::size_t accumulators_ = 4, class Input_ , typename Output_ > | |
| Output_ | pairwise_sum_abstract (const std::size_t num_total, Input_ input, PairwiseSumWorkspace< Output_ > &work, const PairwiseSumOptions &options) |
| template<std::size_t accumulators_ = 4, typename Input_ , typename Output_ > | |
| Output_ | pairwise_sum (const std::size_t num_total, const Input_ *const ptr, PairwiseSumWorkspace< Output_ > &work, const PairwiseSumOptions &options) |
| template<std::size_t accumulators_ = 4, typename Input_ , typename Output_ > | |
| RssResult< Output_ > | rss (const std::size_t num_total, const std::size_t num_non_zero, const Input_ *const ptr, RssWorkspace< Output_ > &work, const RssOptions< Output_ > &options) |
| template<std::size_t accumulators_ = 4, typename Input_ , typename Output_ > | |
| RssResult< Output_ > | rss (const std::size_t num_total, const Input_ *const ptr, RssWorkspace< Output_ > &work, const RssOptions< Output_ > &options) |
| template<typename Output_ = double, typename Input_ , typename Count_ > | |
| void | update_rss (Output_ &mean, Output_ &rss, const Input_ value, const Count_ num_total) |
| template<typename Output_ = double, typename Count_ > | |
| void | update_rss_with_zeros_unsafe (Output_ &mean, Output_ &rss, const Count_ num_zeros, const Count_ num_total) |
| template<typename Output_ = double, typename Count_ > | |
| void | update_rss_with_zeros (Output_ &mean, Output_ &rss, const Count_ num_zeros, const Count_ num_total) |
| template<typename Count_ , typename Float_ > | |
| Float_ | recenter_rss_unsafe (const Count_ num_total, const Float_ old_rss, const Float_ old_mean, const Float_ new_mean) |
| template<typename Count_ , typename Float_ > | |
| Float_ | recenter_rss (const Count_ num_total, const Float_ old_rss, const Float_ old_mean, const Float_ new_mean) |
| template<typename Input_ , class Skip_ > | |
| std::size_t | skip_values (const std::size_t num_total, Input_ *const ptr, Skip_ skip) |
| template<typename Value_ > | |
| constexpr Value_ | nan_if_available_else_zero () |
Quickly compute simple statistics.
| Output_ quickstats::mad | ( | const std::size_t | num_total, |
| Input_ *const | ptr, | ||
| const Input_ | median, | ||
| const MadOptions< Output_ > & | options ) |
Compute the median absolute deviation (MAD) of an array of elements, given its median.
No consideration is given to special values like NaNs in the array. If these are to be skipped, consider using skip_values() before calling this function.
See also scale_mad_to_sd() if the MAD is to be used as an estimate of the standard deviation.
| Output_ | Floating-point type of the output value. |
| Input_ | Numeric type of the input values and median. This is generally expected to be floating-point, though it is also possible to use signed integers as long as their differences do not overflow. |
| num_total | Total number of elements from which to compute a MAD. | |
| [in] | ptr | Pointer to an array of length num_total. On output, the contents will contain the (possibly reordered) absolute deviations from the median. |
| median | Median of the array at ptr, typically computed with median(). For medians that might be infinite, consider using mad_with_infinities() instead. | |
| options | Further options. |
num_total == 0, MadOptions::placeholder is returned. If median is infinite and MadOptions::difference_between_infinites_is_zero == false, MadOptions::placeholder is also returned. | Output_ quickstats::mad | ( | const std::size_t | num_total, |
| const std::size_t | num_non_zero, | ||
| Input_ *const | values, | ||
| const Input_ | median, | ||
| const MadOptions< Output_ > & | options ) |
Compute the median absolute deviation (MAD) of a sparse vector, given its median. This vector is assumed to have num_non_zero structural non-zeros and num_total - num_non_zero zeros.
No consideration is given to special values like NaNs in the array. If these are to be skipped, consider using skip_values() before calling this function.
See also scale_mad_to_sd() if the MAD is to be used as an estimate of the standard deviation.
| Output_ | Floating-point type of the output value. |
| Input_ | Numeric type of the input values and median. This is generally expected to be floating-point, though it is also possible to use signed integers as long as their differences do not overflow. |
| num_total | Total number of elements in the sparse vector. | |
| num_non_zero | Number of structural non-zeros in the sparse vector. This should be no greater than num_total. num_total - num_non_zero is the number of structural zeros. | |
| [in] | values | Pointer to the start of an array of length num_non_zero, containing the values of the structural non-zeros of the sparse vector. On output, the contents will contain the (possibly reordered) absolute deviations from the median. |
| median | Median of the sparse vector at values, typically computed with median(). For medians that might be infinite, consider using mad_with_infinities() instead. | |
| options | Further options. |
num_total == 0, MadOptions::placeholder is returned. If median is infinite and MadOptions::difference_between_infinites_is_zero == false, MadOptions::placeholder is also returned. | Float_ quickstats::scale_mad_to_sd | ( | const Float_ | x | ) |
Scale the median absolute deviation (MAD) so that its expected value for a normal distribution is equal to the standard deviation.
| Float_ | Floating-point type of the MAD. |
| x | MAD, typically computed from mad(). |
x. | Output_ quickstats::median | ( | const std::size_t | num_total, |
| Input_ *const | ptr, | ||
| const MedianOptions< Output_ > & | options ) |
Compute the median of an array of elements.
No consideration is given to special values like NaNs in the array. If these are to be skipped, consider using skip_values() before calling this function.
| num_total | Total number of elements from which to compute a median. | |
| [in] | ptr | Pointer to an array of length num_total. On output, the contents may be reordered. |
| options | Further options. |
| Output_ | Floating-point type of the output value. |
| Input_ | Numeric type of the input values. |
[ptr, ptr + num_total), or MedianOptions::placeholder if num_total == 0. | Output_ quickstats::median | ( | const std::size_t | num_total, |
| const std::size_t | num_non_zero, | ||
| Input_ *const | values, | ||
| const MedianOptions< Output_ > & | options ) |
Compute the median of a sparse vector. This vector is assumed to have num_non_zero structural non-zeros and num_total - num_non_zero zeros.
No consideration is given to special values like NaNs in the array. If these are to be skipped, consider using skip_values() before calling this function.
| num_total | Total number of elements in the sparse vector. | |
| num_non_zero | Number of structural non-zeros in the sparse vector. This should be no greater than num_total. num_total - num_non_zero is the number of structural zeros. | |
| [in] | values | Pointer to the start of an array of length num_non_zero, containing the values of the structural non-zeros of the sparse vector. On output, the contents may be reordered. |
| options | Further options. |
| Output_ | Floating-point type of the output value. |
| Input_ | Numeric type of the input values. |
MedianOptions::placeholder if num_total == 0. | Output_ quickstats::pairwise_sum_abstract | ( | const std::size_t | num_total, |
| Input_ | input, | ||
| PairwiseSumWorkspace< Output_ > & | work, | ||
| const PairwiseSumOptions & | options ) |
Perform pairwise summation on an abstract array of numeric elements. The array is recursively halved until each subarray is no greater than PairwiseSumOptions::max_sum_length. Elements in each subarray are summed in sequence, and the subtotals are then added together to obtain the total sum. Compared to naive summation, this reduces round-off error from floating-point imprecision with minimal loss of performance.
| accumulators_ | Number of accumulators to sum each subarray. This should be positive and should be no greater than PairwiseSumOptions::max_sum_length / 2. It is strongly recommended that this is set to a power of 2. Larger values improve instruction parallelization at the cost of increased binary size and potential register spills. |
| Input_ | Function that accepts a std::size_t and returns an input value (typically numeric). |
| Output_ | Numeric type of the sum. |
| num_total | Total number of observations. |
| input | Function that accesses an abstract array of length num_total. Specifically, it accepts an integer index in [0, num_total) and returns an input value to be summed. Calls to input() may be vectorized so it is assumed that (i) there are no dependencies between calls and (ii) no exceptions are thrown. |
| work | Workspace that can be re-used across multiple pairwise_sum_abstract() calls. |
| options | Further options. |
input(i) values for i in [0, num_total). | Output_ quickstats::pairwise_sum | ( | const std::size_t | num_total, |
| const Input_ *const | ptr, | ||
| PairwiseSumWorkspace< Output_ > & | work, | ||
| const PairwiseSumOptions & | options ) |
Perform pairwise summation on an array, see pairwise_sum_abstract() for more details.
| accumulators_ | Maximum number of elements to sum in sequence, see pairwise_sum_abstract() for details. |
| Input_ | Numeric type of the input data. |
| Output_ | Numeric type of the sum. |
| num_total | Total number of observations. | |
| [in] | ptr | Pointer to an array of length num_total, containing the elements to be summed. |
| work | Workspace that can be re-used across multiple pairwise_sum() calls. | |
| options | Further options. |
[ptr, ptr + num_total). | RssResult< Output_ > quickstats::rss | ( | const std::size_t | num_total, |
| const std::size_t | num_non_zero, | ||
| const Input_ *const | ptr, | ||
| RssWorkspace< Output_ > & | work, | ||
| const RssOptions< Output_ > & | options ) |
Compute the residual sum of squares from a sparse vector. This uses the standard two-pass algorithm with naive accumulation of the sum of squared differences; thus, it is best used with a sufficiently high-precision Output_ like double.
No consideration is given to special values like NaNs in the values of the structural non-zeros. If these are to be skipped, consider using skip_values() before calling this method.
| accumulators_ | Number of accumulators, see pairwise_sum() for details. |
| Input_ | Numeric type of the input values. |
| Output_ | Floating-point type of the output data. |
| num_total | Total number of elements in the sparse vector. | |
| num_non_zero | Number of structural non-zeros in the sparse vector. This should be no greater than num_total. num_total - num_non_zero is the number of structural zeros. | |
| [in] | ptr | Pointer to an array of length num_non_zero, containing the values of the structural non-zeros in the sparse vector. |
| work | Workspace that can be re-used across multiple rss() calls. | |
| options | Further options. |
| RssResult< Output_ > quickstats::rss | ( | const std::size_t | num_total, |
| const Input_ *const | ptr, | ||
| RssWorkspace< Output_ > & | work, | ||
| const RssOptions< Output_ > & | options ) |
Compute the residual sum of squares from the mean from a dense array. This uses the standard two-pass algorithm with naive accumulation of the sum of squared differences; thus, it is best used with a sufficiently high-precision Output_ like double.
No consideration is given to special values like NaNs in the dense array. If these are to be skipped, consider using skip_values() before calling this method.
| accumulators_ | Number of accumulators, see pairwise_sum() for details. |
| Input_ | Numeric type of the input values. |
| Output_ | Floating-point type of the output data. |
| num_total | Total number of elements in the array. | |
| [in] | ptr | Pointer to an array of length num_total. |
| work | Workspace that can be re-used across multiple rss() calls. | |
| options | Further options. |
| void quickstats::update_rss | ( | Output_ & | mean, |
| Output_ & | rss, | ||
| const Input_ | value, | ||
| const Count_ | num_total ) |
Update the mean and RSS by adding a new value using Welford's method.
| Output_ | Floating-point type of the output statistics. |
| Input_ | Numeric type of the inut value. |
| Count_ | Integer type of the number of values. |
| mean | On input, the mean of previous values. If no previous values were provided, this should be set to zero. On output, the updated mean after including the latest value. |
| rss | On input, the RSS of previous values. If no previous values were provided, this should be set to zero. On output, the updated RSS after including the latest value. |
| value | New value to update the mean/RSS. |
| num_total | Number of values used to compute the updated mean/RSS, after adding the latest value. This should always be positive. |
| void quickstats::update_rss_with_zeros_unsafe | ( | Output_ & | mean, |
| Output_ & | rss, | ||
| const Count_ | num_zeros, | ||
| const Count_ | num_total ) |
Update the mean and RSS by adding any number of zeros using Welford's method. This assumes that num_total > 0; if this cannot be guaranteed, use update_rss_with_zeros() instead.
| Output_ | Floating-point type of the output statistics. |
| Count_ | Integer type of the number of values. |
| mean | On input, the mean of previous values. If no previous values were provided, this should be set to zero. On output, the updated mean after including the zeros. |
| rss | On input, the RSS of previous values. If no previous values were provided, this should be set to zero. On output, the updated RSS after including the zeros. |
| num_zeros | Number of zeros to be added. This may be zero. |
| num_total | Number of values used to compute the updated mean/RSS, after adding the specified number of zeros. This should be positive and no less than num_total. |
| void quickstats::update_rss_with_zeros | ( | Output_ & | mean, |
| Output_ & | rss, | ||
| const Count_ | num_zeros, | ||
| const Count_ | num_total ) |
Update the mean and RSS with any number of zeros using Welford's method. This is a slightly slower version of update_rss_with_zeros_unsafe() that handles num_total == 0.
| Output_ | Floating-point type of the output statistics. |
| Count_ | Integer type of the number of values. |
| mean | On input, the mean of previous values. If no previous values were provided, this should be set to zero. On output, the updated mean after including the zeros. |
| rss | On input, the RSS of previous values. If no previous values were provided, this should be set to zero. On output, the updated RSS after including the zeros. |
| num_zeros | Number of zero values to be added. This may be zero. |
| num_total | Number of values used to compute the updated mean/RSS, after adding any number of zeros. This may be zero but should be no less than num_total. |
| Float_ quickstats::recenter_rss_unsafe | ( | const Count_ | num_total, |
| const Float_ | old_rss, | ||
| const Float_ | old_mean, | ||
| const Float_ | new_mean ) |
Recenter the residual sum of squares, i.e., sum of squares from a different mean. This is typically used to combine RSS values from different subsets of the data, e.g., when splitting calculations across cores. In such cases, the global mean across the entire dataset should be used as the new mean, and the sum of the recentered RSS values will be the RSS of the entire dataset. (This approach is more numerically stable than computing the sum of squared observations and then computing the difference with the squared mean.)
This function is considered "unsafe" as it assumes that old_mean is zero when num_total == 0. In many cases, old_mean will be NaN when num_total == 0 due to division by zero. If num_total > 0 or old_mean == 0 cannot be guaranteed, consider using recenter_rss() instead.
| Count_ | Integer type of the number of values. |
| Float_ | Floating-point type of the various statistics. |
| num_total | Total number of values used to compute the RSS. |
| old_rss | The old value of the RSS. |
| old_mean | The old mean used to compute the RSS. |
| new_mean | The new mean. |
| Float_ quickstats::recenter_rss | ( | const Count_ | num_total, |
| const Float_ | old_rss, | ||
| const Float_ | old_mean, | ||
| const Float_ | new_mean ) |
Recenter the residual sum of squares, i.e., sum of squares from a different mean. This is a safer version of recenter_rss_unsafe() that correctly handles num_total == 0 and an old_mean of NaN, at the cost of some performance.
| Count_ | Integer type of the number of values. |
| Float_ | Floating-point type of the various statistics. |
| num_total | Total number of values used to compute the RSS. This should be non-negative. |
| old_rss | The old value of the RSS. |
| old_mean | The old mean used to compute the RSS. If num_total == 0, this is ignored and may be set to NaN; this NaN will not propagate to the output. |
| new_mean | The new mean. |
old_rss (which should be zero) if num_total == 0. | std::size_t quickstats::skip_values | ( | const std::size_t | num_total, |
| Input_ *const | ptr, | ||
| Skip_ | skip ) |
| Input_ | Numeric type of the input. |
| Skip_ | Function defining what elements to skip. |
| num_total | Total number of elements. | |
| [in,out] | ptr | Pointer to an array of length num_total, containing the input elements. On output, the first \(n\) entries will contain the unskipped elements, in the same order as provided in the input. |
| skip | Function that accepts the index of an element in ptr (as a std::size_t) and the value of the element (as an Input_), and returns a boolean indicating whether this element should be skipped. |
|
constexpr |
| Value_ | Some numeric type. |
Value_, otherwise zero.