A Sampling-Based Computational Strategy for the Representation of Epistemic Uncertainty in Model Predictions with Evidence Theory

Summary

This report presents a sampling-based computational strategy for representing epistemic uncertainty in model predictions using evidence theory (Dempster-Shafer theory). The approach addresses the high computational cost and dimensionality challenges of evidence propagation through complex models by utilizing Latin hypercube sampling, sensitivity analysis, and nonparametric response surface approximations (MARS). The methodology is illustrated on a thermal safety reliability problem involving a weak link/strong link system.

Cover Page

SANDIA REPORT SAND2006-5557 Unlimited Release Printed October 2006

A Sampling-Based Computational Strategy for the Representation of Epistemic Uncertainty in Model Predictions with Evidence Theory

J.C. Helton, J.D. Johnson, W.L. Oberkampf, C.B. Storlie

Prepared by Sandia National Laboratories Albuquerque, New Mexico 87185 and Livermore, California 94550

Sandia is a multiprogram laboratory operated by Sandia Corporation, a Lockheed Martin Company, for the United States Department of Energy’s National Nuclear Security Administration under Contract DE-AC04-94AL85000.

Approved for public release; further dissemination unlimited.

Notice and Availability

Issued by Sandia National Laboratories, operated for the United States Department of Energy by Sandia Corporation.

NOTICE: This report was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government, nor any agency thereof, nor any of their employees, nor any of their contractors, subcontractors, or their employees, make any warranty, express or implied, or assume any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represent that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise, does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government, any agency thereof, or any of their contractors or subcontractors. The views and opinions expressed herein do not necessarily state or reflect those of the United States Government, any agency thereof, or any of their contractors.

Printed in the United States of America. This report has been reproduced directly from the best available copy.

Available to DOE and DOE contractors from: U.S. Department of Energy Office of Scientific and Technical Information P.O. Box 62 Oak Ridge, TN 37831 Telephone: (865)576-8401 Facsimile: (865)576-5728 E-Mail: [email protected] Online ordering: http://www.doc.gov/bridge

Available to the public from: U.S. Department of Commerce National Technical Information Service 5285 Port Royal Rd. Springfield, VA 22161 Telephone: (800)553-6847 Facsimile: (703)605-6900 E-Mail: [email protected] Online ordering: http://www.ntis.gov/ordering.htm

Title and Abstract

SAND2006-5557 Unlimited Release Printed October 2006

A Sampling-Based Computational Strategy for the Representation of Epistemic Uncertainty in Model Predictions with Evidence Theory

J.C. Helton,a J.D. Johnson,b W.L. Oberkampf,c C.B. Storlied aDepartment of Mathematics and Statistics, Arizona State University, Tempe, AZ 85287-1804 USA bProStat, Mesa, AZ 85204-5326 USA cSandia National Laboratories, Albuquerque, NM 87185-0828 USA dDepartment of Statistics, North Carolina State University, Raleigh, NC 27695 USA

Abstract Evidence theory provides an alternative to probability theory for the representation of epistemic uncertainty in model predictions that derives from epistemic uncertainty in model inputs, where the descriptor epistemic is used to indicate uncertainty that derives from a lack of knowledge with respect to the appropriate values to use for various inputs to the model. The potential benefit, and hence appeal, of evidence theory is that it allows a less restrictive specification of uncertainty than is possible within the axiomatic structure on which probability theory is based. Unfortunately, the propagation of an evidence theory representation for uncertainty through a model is more computationally demanding than the propagation of a probabilistic representation for uncertainty, with this difficulty constituting a serious obstacle to the use of evidence theory in the representation of uncertainty in predictions obtained from computationally intensive models. This presentation describes and illustrates a sampling-based computational strategy for the representation of epistemic uncertainty in model predictions with evidence theory. Preliminary trials indicate that the presented strategy can be used to propagate uncertainty representations based on evidence theory in analysis situations where naïve sampling-based (i.e., unsophisticated Monte Carlo) procedures are impracticable due to computational cost.

Key Words: Dempster-Shafer theory, Epistemic uncertainty, Evidence theory, Monte Carlo, Numerical uncertainty propagation, Sensitivity analysis, Uncertainty analysis

Acknowledgments

Acknowledgments

Work performed at Sandia National Laboratories (SNL), which is a multiprogram laboratory operated by Sandia Corporation, a Lockheed Martin Company, for the U.S. Department of Energy’s National Nuclear Security Administration under Contract No. DE-AC04-04AL85000. Review at SNL provided by L. Swiler and C. Sallaberry.

Editorial support provided by F. Puffer, J. Ripple, and K. Best of Tech Reps, a division of Ktech Corporation.

Contents and Lists of Figures and Tables

Contents

Acknowledgments … 4

  1. Introduction … 7
  2. Evidence Theory … 9
  3. Computational Strategy for Estimating CBF, CCBF, CPF and CCPF for y … 17
  4. Example for Illustration … 19
  5. Example Results … 23 5.1 Step 1: Define Sampling Distribution … 23 5.2 Step 2: Generate Latin Hypercube Sample … 23 5.3 Step 3: Propagate Sample Through Model … 23 5.4 Step 4: Perform Sensitivity Analysis … 23 5.5 Step 5: Develop Response Surface Approximation … 25 5.6 Step 6: Approximate y for Large Random Sample … 28 5.7 Step 7: Approximate Evidence Space Results … 30
  6. Discussion … 37
  7. References … 39

Figures Fig. 1. Form of the CPF, CDF, CBF, CCPF, CCDF and CCBF that results for each variable from the uncertainty information in Table 2 with variable range normalized to [0, 1] … 22 Fig. 2. Time-dependent curves for WLs and SLs obtained with Latin hypercube sample with 100 curves from sample of size 200 shown … 24 Fig. 3. Values for pF obtained with Latin hypercube sample of size 200 shown as a CCDF … 24 Fig. 4. Comparison of observed values and MARS response surface predicted values: (a) WL1T75, and (b) log(pF) … 29 Fig. 5. Stepwise construction of CCBFs and CCPFs for WL1T75 with c61, c2 and c1: (a) Construction with MARS response surface approximation, and (b) Construction with predicted values … 31 Fig. 6. Stepwise construction of CCBFs and CCPFs with MARS response surface approximations for WL1T25, SL1T25 and SL1T75 … 32 Fig. 7. Stepwise construction of CCBFs and CCPFs for pF with c71, c2, c1, c72 and c9 … 33 Fig. 8. Simplification of an evidence space with each horizontal line corresponding to a focal element with a BPA of 0.2 in a new evidence space … 34 Fig. 9. Stepwise construction of CCBFs and CCPFs for pF with the evidence spaces redefined to have 5 focal elements … 35 Fig. 10. Stepwise construction of CCBFs and CCPFs for pF with alternative ordering … 36

Tables Table 1. Uncertain Variables and Associated Uncertainty Ranges Considered in Example Uncertainty Analyses … 21 Table 2. Illustrative Specification of Uncertainty Information Used in Example Uncertainty Analyses with Probability Theory and Evidence Theory for Variables in Table 1 … 21 Table 3. Basic Probability Assignments (BPAs) for a Variable on the Interval [a, b] Derived from the Information in Table 2 … 22 Table 4. Sensitivity Analysis Based on Stepwise Rank Regression for Mapping [xi, yi], y = 1, 2, …, nS = 200 … 26 Table 5. Summaries of Stepwise Construction of Response Surface Approximations of log(pF) … 28 Table 6. Summaries of Stepwise Construction of Response Surface Approximations to WL1T25, WL5T75, SL1T25 and SL1T75 with the MARS Procedure … 29

1. Introduction

  1. Introduction

An appropriate representation of the uncertainty in analysis outcomes is an essential part of any complete analysis. Specifically, an analysis that is intended to provide insights into the behavior of a system or the basis for decisions must provide an assessment of the uncertainty associated with its outcomes. Without such an assessment, neither insights drawn from the analysis nor decisions based on it are adequately informed and supported.

Analyses of the behavior of complex systems typically involve two types of uncertainty: aleatory and epistemic. Aleatory uncertainty arises from what is considered to be an inherent randomness in the behavior of the system under study. For example, in a risk assessment for a chemical plant, the weather conditions at the time of an accident are usually considered to be an aleatory uncertainty. Alternatives to the descriptor aleatory include stochastic, variability, irreducible and type A. Epistemic uncertainty arises from a lack of knowledge about a quantity that is assumed to have a fixed value in the context of a particular analysis. For example, the pressure at which a specific reactor containment will fail is presumably fixed but certainly unknown and is thus an epistemic uncertainty. Alternatives to the descriptor epistemic include subjective, state of knowledge, reducible and type B. As an example, probabilistic risk assessments for nuclear power plants are typically designed to maintain a separation between aleatory uncertainty and epistemic uncertainty.

Probability has traditionally been employed as the mathematical structure used to represent both aleatory uncertainty and epistemic uncertainty. With this usage, an analysis maintaining a separation of aleatory uncertainty and epistemic uncertainty involves two probability spaces: one probability space characterizing aleatory uncertainty and one probability space characterizing epistemic uncertainty. This dual usage of probability can be traced back to at least the beginnings of the formal development of probability theory in the late seventeenth century. However, many individuals have reservations about the use of probability to represent epistemic uncertainty when there is limited information available on which to base a fully structured development of probability. In particular, the concern is that the definition of a full probabilistic description of uncertainty entails an implication of a higher resolution of knowledge than is really present.

Evidence theory provides an alternative to probability theory for the representation of epistemic uncertainty in model predictions that derives from epistemic uncertainty in model inputs. The potential benefit, and hence appeal, of evidence theory is that it allows a less restrictive specification of uncertainty than is possible within the axiomatic structure on which probability theory is based. Unfortunately, the propagation of an evidence theory representation for uncertainty through a model is more computationally demanding than the propagation of a probabilistic representation for uncertainty, with this difficulty constituting a serious obstacle to the use of evidence theory in the representation of uncertainty in predictions obtained from computationally intensive models. This presentation describes and illustrates a sampling-based computational strategy for the representation of epistemic uncertainty in model predictions with evidence theory. Preliminary trials indicate that the presented strategy can be used to propagate uncertainty representations based on evidence theory in analysis situations where naïve sampling-based (i.e., unsophisticated Monte Carlo) procedures are impracticable due to computational cost.

This presentation is organized as follows. First, an overview of evidence theory is given (Sect. 2). Then, the numerical procedure for the construction of evidence theory results is described (Sect. 3). This description is followed by the introduction of the illustrative example involving a weak link (WL)/strong link (SL) system (Sect. 4) and the presentation of results obtained with evidence theory in the analysis of this system (Sect. 5). Finally, the presentation ends with a concluding discussion (Sect. 6).

2. Evidence Theory

  1. Evidence Theory

An analysis can be conceptually represented in the functional form y = f(x), (2.1) where x = [x1, x2, …, xnX] (2.2) is a vector of analysis inputs, and y = [y1, y2, …, ynY] (2.3) is a vector of analysis results, and f is a function that maps x into y. In practice, f can be quite complex and could involve the solution of a system of nonlinear partial differential equations or the operation of a sequence of linked models.

Probability theory provides the mathematical structure that has been traditionally used to characterize epistemic uncertainty in x through a sequence of distributions D1, D2, …, DnX, (2.4) which give rise to a probability space (Xp, Xp, mpx). The uncertainty in x gives rise to uncertainty in y. For a single real-valued component y = f(x) (2.5), the CDF and CCDF are defined by: prob(y~ <= y) = integral_{Xp} delta[f(x) <= y] dx(x) dX = sum_{i=1}^{nS} delta[f(xi) <= y] / nS (2.6) prob(y~ > y) = integral_{Xp} delta_bar[f(x) > y] dx(x) dX = sum_{i=1}^{nS} delta_bar[f(xi) > y] / nS (2.7) where delta[f(x) <= y] = 1 if f(x) <= y, and 0 if f(x) > y (2.8), and delta_bar = 1 - delta (2.9).

A sample xi = [xi1, xi2, …, xi,nX], i = 1, 2, …, nS (2.10) is typically generated via Latin hypercube sampling.

Alternatives to probability theory include interval analysis, possibility theory, fuzzy set theory, and evidence theory (Dempster-Shafer theory). An evidence space for x is a triple (XE, XE, mEX) where XE is the sample space, XE is a set of subsets of XE, and mEX is a basic probability assignment (BPA) satisfying: mEX(U) > 0 if U in XE and U != empty set (2.16) mEX(U) = 0 if U subset XE and U not in XE (2.17) sum_{U in XE} mEX(U) = 1 (2.18)

Measures of uncertainty in evidence theory are belief (Bel) and plausibility (Pl): BelX(U) = sum_{V subset U} mEX(V) (2.20) PlX(U) = sum_{V cap U != empty set} mEX(V) (2.21) Relationships include: BelX(U) + BelX(U^c) <= 1 (2.22) PlX(U) + PlX(U^c) >= 1 (2.23) PlX(U) + BelX(U^c) = 1 (2.24)

For multidimensional x with independent components: XE = {U : U = U1r x U2s x … x UnX,t} (2.27) mEX(U) = mE1(U1r) * mE2(U2s) * … * mE,nX(UnX,t) (2.28) The number of focal elements is n = prod_{j=1}^{nX} n(j) (2.29), which grows exponentially with dimensionality.

Cumulative belief functions (CBF), complementary cumulative belief functions (CCBF), cumulative plausibility functions (CPF), and complementary cumulative plausibility functions (CCPF) characterize the epistemic uncertainty in y.

3. Computational Strategy for Estimating CBF, CCBF, CPF and CCPF for y

  1. Computational Strategy for Estimating CBF, CCBF, CPF and CCPF for y

The proposed computational strategy involves the following initial steps:

Step 1. Define a sampling distribution for x based on the specified evidence theory structure for the uncertain model inputs, using density function dj(xj) = sum_{k=1}^{n(j)} delta(xj | Ujk) mEj(Ujk) / (bjk - ajk).

Step 2. Generate a Latin hypercube sample from the uncertain inputs with the sampling distribution defined in Step 1 (sample size nS).

Step 3. Propagate the sample generated in Step 2 through the model to obtain values for all model results of interest, producing mapping [xi, yi].

The following additional steps are performed individually for each model result y of interest:

Step 4. Perform a sensitivity analysis to identify which uncertain model inputs (x1, x2, …, xr) are significant contributors to the uncertainty in y.

Step 5. Use regression procedures (linear regression, LOESS, GAMs, PP_REG, or MARS) to develop a response surface approximation to y as a function of the significant variables.

Step 6. Generate a large random sample (e.g., nS = 10^6) from the uncertain inputs and evaluate y using the response surface approximation.

Step 7. Perform a sequential construction of the CBF, CCBF, CPF, and CCPF for y with the response surface results, adding variables sequentially until the functions converge or all important variables are incorporated.

4. Example for Illustration

  1. Example for Illustration

The example involves a system with two weak links (WLs) and two strong links (SLs) in a fire accident environment, adapted from a competing risk problem. The failure of both SLs before the failure of either WL represents the undesirable event: Probability of Loss of Assured Safety (PLOAS, pF).

Equations: TMPWLj(t) = c1 + [c2 + c3j * exp(-c4j * t) * sin(c5j * t)] * tanh(c6j * t) (4.7) TMPSLk(t) = c1 + c2 * tanh[c62 * (1 + c7k) * t] (4.8) Failure temperatures follow normal distributions with means and standard deviations: fWLj(TWL) = (1 / (c9 * sqrt(2pi))) * exp(-(TWL - c8)^2 / (2 * c9^2)) (4.9) fSLk(TSL) = (1 / (c11 * sqrt(2pi))) * exp(-(TSL - c10)^2 / (2 * c11^2)) (4.10)

Table 1. Uncertain Variables and Ranges:

  • c1: Initial temperature [-30, 40 °C]
  • c2: Steady-state temperature increase [800, 1000 °C]
  • c31, c32: Peak amplitude of transient [-2600, -100 °C]
  • c41, c42: Thermal heating time constant [0.2, 0.4 min^-1]
  • c51, c52: Frequency response of transient [0.1, 0.2 min^-1]
  • c61, c62: Rate time constant [0.01-0.015 min^-1] and [0.021-0.025 min^-1]
  • c71, c72: Rapid heating factor in SLs [0.3, 0.5] and [0.6, 2.0]
  • c8, c9: WL failure temp mean [285, 315 °C] and std dev [4, 12 °C]
  • c10, c11: SL failure temp mean [560, 580 °C] and std dev [15, 35 °C]

Four expert opinions are elicited for each variable (Table 2), combined equally into 13 focal elements with assigned BPAs (Table 3).

5. Example Results

  1. Example Results

Results of applying the 7-step strategy:

  • A Latin hypercube sample of size nS = 200 was generated and propagated through the model.
  • Stepwise rank regression identified the dominant inputs (Table 4):
    • For WL1T25: c61, c1, c2 (R^2 = 0.96)
    • For WL1T75: c61, c2, c1 (R^2 = 0.96)
    • For SL1T25: c2, c62, c71, c1, c31 (R^2 = 0.94)
    • For SL1T75: c2, c1, c71 (R^2 = 0.97)
    • For log(pF): c71, c2, c1, c72, c8, c9, c10, c11, c42, c32, c62, c52 (R^2 = 0.93)
  • Stepwise MARS models were constructed (Tables 5 and 6) and validated using ‘leave-one-out’ cross-validation.
  • A sample of size nS = 10^6 was generated to construct CCBF and CCPF bounds sequentially.
  • Simplified 5-focal-element representations (Fig. 8) enabled efficient stepwise inclusion of additional variables, yielding tighter and well-converged bounds for log(pF) (Figs. 9 and 10).

6. Discussion

  1. Discussion

Evidence theory is a promising alternative to probability theory for the representation of epistemic uncertainty when limited information is available, providing a less structured framework.

Two interpretations exist: (1) specification of an incompletely defined probability space (belief is minimum probability, plausibility is maximum probability), and (2) a structure for reasoning under uncertainty (belief measures supportive information, plausibility measures absence of contradictory information).

The sampling-based computational strategy using Latin hypercube sampling and nonparametric regression (MARS) enables propagation through computationally demanding models where standard Monte Carlo fails due to the curse of dimensionality. The stepwise process also generates useful sensitivity analysis insights within the evidence theory framework.

7. References

  1. References

Contains 158 cited references spanning verification and validation, uncertainty quantification, sensitivity analysis, evidence theory, Monte Carlo methods, and nonparametric regression techniques.

Distribution

Distribution List Includes external distribution to academic institutions, national laboratories (LANL, LLNL, ANL), DOE, NRC, and international researchers, as well as internal Sandia National Laboratories distribution.