MindMap-ReliabilityHandbook
April 23, 2017 Leave a comment
MindMap-ReliabilityHandbook
- Basic Reliability Concept
- Goal
- MTBF=Long
- MTTR=Short
- MTR=Short
- MTBI=Less
- Reliability Program Plan
- Immediate issues
- Reactive reliability engineering
- Overall process
- Reliability engineering metrics with examples
- Fault vs. failure
- Pareto plot
- Uptime and Availability
- MTBF, MTBA, MTBI
- MTTR, MTR
- Design For Reliability Process
- 1. Initial Design
- 2. Developmental Testing: Failure Mode Discovery
- 2.1. Root Cause Analysis
- 2.2. Development of Corrective Actions
- 2.3. Corrective Action Review and Approval
- 2.4. Assignment of Fix Effectiveness Factors
- 2.5. Verification of Corrective Actions
- 3. Final Design: Meets Requirement
- 4. Demonstration Testing
- Reliability Life Cycle
- 1. Reliability Specification
- 2. Reliability Design
- 3. Reliability Verification
- 4. Reliability Maintenance
- 5. Reliability Audit
- NOTE: Cannot effectively “Test-in” reliability on Step 5. Too late.
- Phases of Reliability Process
- Phase1: Set Reliability Goals
- Establish reliability goals, such as “X” Reliability at “Y%” confidence after “N” operations or time on “Z” performance parameter for the product.
Risk Mitigation planning – QFD / FMEA
Reliability Metrics: generate quantitative, measurable failure definitions
Operating environments: define the operating environment that encourages a failure event.
In General, 65% of Life Cycle Cost is Fixed at This Stage
- Establish reliability goals, such as “X” Reliability at “Y%” confidence after “N” operations or time on “Z” performance parameter for the product.
- Phase2: Develop System Models
- Identify the reliability system elements (subsystem & components)
Construct a block diagram for each system element
Identify critical components and failure potentials within each system element (i.e. the x’s)
Allocate the reliability metrics to the lowest level
Benchmark the current design with this model.
Determine acceleration or deceleration (mitigation) factors (i.e.temperature, loads, cycles, etc.)
Identify potential improvements to reliability, cost, features based on the analysis. Assess tradeoffs associated with each.
System Modeling
- Identify the reliability system elements (subsystem & components)
- Phase3: Design Product for Reliability
- Apply robust design simulation tools
- Determine Key Control Parameters (KCP’s) and Key Noise Parameters(KNP’s).
- Map reliability requirements to critical performance or process elements
- Use tools analytically & experimentally on the KCP’s & KNP’s
- Conduct life predictions through reliability modeling.
- Develop risk mitigation plans, including analysis of service and warranty issues.
- In General, 85% of Life Cycle Cost is Fixed at This Stage
- Phase4: Perform Reliability Verification
- Develop reliability test strategies
- Internal
- – short term
- – long term
- External (Field tests, Beta sites, etc.)
- – short term
- – long term
- Accelerated Testing
- – HALT/ HASS
- – failure mode testing ( Address FMEA/FTA concerns)
- Compliance (Safety) Testing
- – agencies
- – customers
- Internal
- In General, 95% of Life Cycle Cost is Fixed at This Stage
- Develop reliability test strategies
- Phase5: Implement Production and Field Reliability Systems
- Establish Audit Program
- Process verification
- Product verification
- Establish FRACAS a system
- Correlate field data with all test results
- Identify means for collecting continuous data (not just failure data).
- In General, 98% of Life Cycle Cost is Fixed at This Stage
- NOTE: Design for Reliability Should be Integrated into the Design Process
- Phase1: Set Reliability Goals
- Legacy Product Reliability Growth Process
- Phase1: Review Historical Data
- Review historical endurance data
- Review field RMAs
- Review customer environment & application
- Phase2: Analyze Field & In-house Endurance Test Data
- Develop product Fault Tree Analysis
- Identify and pareto observed failure modes
- Phase3: Develop Reliability Profile & Goals
- Develop P-Diagram and System Block Diagram
- Generate Reliability plots for operational endurance
- Allocate reliability goals to key subsystems
- Identify reliability gaps between existing product & goals for each subsystem
- Phase4: Develop & Execute Reliability Growth Plan
- Determine root cause for all identified failures
- Redesign process or parts to address failure mode pareto
- Validate reliability improvement through accelerated life testing & field betas
- Allocate key part and subassembly CTQs to suppliers
- Phase5: Institute Reliability Validation Program
- Implement design and process firewalls & sensors to hold design robustness
- Develop and implement long-term reliability validation audit
- Phase1: Review Historical Data
- 12 Steps: Creating FMECA
- 1. Learn about the system
- Know about the system
- Understand the system
- Identify key characteristics of the system
- Identify key subsystem, components
- 2. Set the level of the analysis
- be aware that FMECA is a bottom up technique of problem solving
- define the scope of the system to study
- ask the following questions:
- how components can fail?
- what are the causes of failure
- 3. Describe the desired functions
- Know the function of the system
- Know the purpose of the system
- 4. Review and rationalise the functions
- 5. List all potential failure modes
- 6. Review and rationalise failure modes against functions
- 7. Describe the potential effects of each failure
- 8. Describe the potential causes of each failure
- 9. Describe the current controls
- 10.Assess criticality
- 11.Take corrective action
- 12.Standardise the actions taken
- 1. Learn about the system
- Goal
- Reliability Definitions
- Fault: anything that has gone wrong
- Failure: an equipment problem
- All failures are faults
- Examples:
- 1. If a transport system stops due to particles that are normal to the process, then it is a failure (and a fault).
2. a wrench left inside the equipment, then it’s a fault but not a failure.
- 1. If a transport system stops due to particles that are normal to the process, then it is a failure (and a fault).
- Examples:
- All time: either “uptime” or “downtime”
- Uptime: either operating or idle time
- Uptime (hours) « availability (%)
- Downtime: either PM, or Unscheduled Maintenance (Repairs)
- MTTR (mean time to repair) applies to PM and to UM
- Equipment Availability
- The probability that the equipment will be in a condition to perform its intended function
- Total Time
- Operation Time
- Uptime
- Manufacturing Time
- Productive Time
- Standby Time
- Manufacturing Time
- Downtime
- Unscheduled Down Time
- Scheduled Down Time
- Uptime
- Operation Time
- Assist: an unplanned interruption where
- Externally resumed (human operator or host computer), and
- No replacement of parts, other than specified consumables, and
- No further variation from specifications of equipment operation
- Machine Availability
- Specification: Availability >85%
- Typical performance: 90%
- Failure: unplanned interruption that is not an assist
- # of interrupts = # of assists + # of failures
- MTBF: Mean Time Between Failures [hours]
- MTBA: Mean Time Between Assist
- MTBI: Mean Time Between Interrups
- MTBF = Interval / (number of failures)
- MTBA = Interval / (number of assists)
- MTBI = Interval / (number of interrupts)
- MTTR: Mean Time To Repair
- The average elapsed time (not person hours) to correct a failure and return the equipment to a condition where it can perform its intended function, including equipment test time and process test time (but not maintenance delay).
- MTR: Mean Time to Restore
- Includes maintenance delays
- Reliability Engineering Metrics
- Fault vs. failure: all failures are faults
- Pareto plot: location and function, sample size of several
- Uptime and Availability: time is up or down
- MTBF, MTBA, MTBI: I = F + A
- MTTR, MTR: working time vs. clock time
- DFMEA: Design Failure Mode and Effect Analysis
- DFR: Design for Reliability
- HALT: Highly Accelerated Life Testing
- Stress levels are increased as far as failures occur.
Therefore this method is considered as a qualitative accelerated test method.
The purpose is to support DFMEA for complex design where traditional brainstorming is not sufficient and there is a risk that existing design guidelines do not cover new technology and/or features under design verification
- Stress levels are increased as far as failures occur.
- HASS: Highly Accelerated Stress Testing
- PoF and HALT techniques are employed to expedite the time between any potential failures and corrective actions
- OPE: Operational Test and Evaluation
- PoF: Physics of Failure
- Physics of Failure (PoF) analyses provide a science-based approach to reliability that utilizes modeling and simulation to design-in reliability. The analyses use computer-aided design tools to model the root causes of failures such as fatigue, fracture, wear and corrosion. The basic approach involves the following:
- 1. Identifying potential failure mechanisms (chemical, electrical, physical, mechanical, structural or thermal processes leading to failures); failure sites; and failure modes even before formal testing is complete.
2. Identifying the appropriate failure models and their input parameters, such as material characteristics, damage properties, manufacturing flaw and defects, etc.
3. Determining where variability exists with respect to each design parameter.
4. Computing the effective reliability function.
5. Accepting the current design or proposed design with new corrective action, if the estimated reliability exceeds the goal over the simulated time period.
6. Proactively incorporating reliability into the design process by establishing a scientific basis for evaluating new materials, structures, and electronics technologies.
7. Using generic failure models, where appropriate, that are as effective for new materials and structures as they are for existing designs.
8. Encouraging innovative, cost-effective design through the use of realistic reliability assessment
- 1. Identifying potential failure mechanisms (chemical, electrical, physical, mechanical, structural or thermal processes leading to failures); failure sites; and failure modes even before formal testing is complete.
- Physics of Failure (PoF) analyses provide a science-based approach to reliability that utilizes modeling and simulation to design-in reliability. The analyses use computer-aided design tools to model the root causes of failures such as fatigue, fracture, wear and corrosion. The basic approach involves the following:
- LIRP: Low Rate Initial Production
- Level of Maturity
- RCA: Root Cause Analysis
- purpose of the root cause analysis is to find the exact reason why the item failed
- QALT: Quantitative Accelerated Lifetime Testing
- RBD: Reliability Block Diagram
- models are often used to describe the relationship between the system components in order to determine the reliability of the system as a whole
- IFR: Instantaneous Failure Rate
- hazard rate
- V&V Verification and Validation
- FMEA: Failure Mode and Effect Analysis
- 1. a failure modes and effects analysis (FMEA) is a procedure in operations management for analysis of potential failure modes within a system,
2. conducted to classify the failure modes by severity or determine of the effect of failures on the system.
3. provides an analytical approach when dealing with potential failure modes and their associated causes.
4. identification of potential failure modes and causes provides engineers with the information needed to alter the development/manufacturing process in order to optimize the tradeoffs between
- 1. safety
- 2. cost
- 3. performance
- 4. quality
- 5. reliability.
5. provides an easy tool to determine which risk has the greatest concern, and thereby what action is needed to prevent a problem before it arises.
6. provides development of robust process specifications to ensure the end product will meet the predefined reliability requirements
- 1. a failure modes and effects analysis (FMEA) is a procedure in operations management for analysis of potential failure modes within a system,
- FMECA: Failure Mode and Effect Criticality Analysis
- 1. is an extension of the FMEA.
2. FMECA is a reliability evaluation and design review technique that examines the potential failure modes within a system or lower piece-part/component level, in order to determine the effects of component level failures on total equipment or system level performance.
3. Each hardware or software failure mode is classified according to its impact on system operation success and personnel safety.
4. FMECA uses inductive logic (a process of finding explanations) on a “bottom up” system hierarchy and traces up through the system hierarchy to determine the end effect on system performance.
5. Maximum benefits are seen when FMECA is conducted early in the design cycle rather than after the design is finalized.
Inputs. The FMEA/FMECA requires the following basic inputs:
- 1. Item(s)
- 2. Function(s)
- 3. Failure (s)
- 4. Effect(s) of Failure
- 5. Cause(s) of Failure
- 6. Current Control(s)
- 7. Recommended Action(s)
RPN: Risk Priority Number
- 1. Rate the severity of each effect of failure.
2. Rate the likelihood of occurrence for each cause of failure.
3. Rate the likelihood of prior detection for each cause of failure (i.e. the likelihood of detecting the problem before it reaches the end user or customer).
4. Calculate the RPN by obtaining the product of the three ratings:
RPN = Severity x Occurrence x Detection
FMECA Process Checlist
- 1. Is the system definition/description provided compatible with the system specification? Inaccurate documentation may result in incorrect analysis conclusions
2. Are ground rules clearly stated? These include approach, failure definition, acceptable degradation limits, level of analysis, clear description of failure causes, etc.
3. Are block diagrams provided showing functional dependencies at all equipment piece-part levels? This diagram should graphically show what items (parts, circuit cards, subsystems, etc.) are required for the successful operation of the next higher assembly
4. Does the failure effect analysis start at the lowest hardware level and systematically work to higher piece part levels? The analysis should start at the lowest level appropriate to make design changes (e.g. part, circuit card, subsystem, etc.)
5. Are failure mode data sources fully described? Specifically identify data sources, including relevant data from similar systems
6. Are detailed FMECA worksheets provided? Do the worksheets clearly track from lower to higher hardware levels? Do the worksheets clearly correspond to the block diagrams? Do the worksheets provide an adequate scope of analysis? Worksheets should provide an item name, piece-part code/number, item function, list of item failure modes, effect on next higher assembly and system for each failuremode, and a critically ranking. In addition, worksheets should account for multiple failure piece-part levels for catastrophic and critical failures.
7. Are failure severity classes provided? Are specific failure definitions established? Typical classes are (see MIL-STD-882D for details): Catastrophic (life/death) Critical (mission loss) Marginal (mission degradation) Minor (maintenance/repair)
8. Are results timely? Analysis must be performed early during the design phase not after the fact
9. Are results clearly summarized and comprehensive recommendations provided? Actions for risk reduction of single point failures, critical items, areas needing built-in test (BIT), etc.
10. Are the results being communicated to enhance other program decisions? Critical parts, reliability prediction, derating, fault tolerance
Quantitative Criticality Analysis Method.
- Define the reliability/unreliability for each item, at a given operating time.
Identify the portion of the item‟s unreliability that can be attributed to each potential failure mode.
Rate the probability of loss (or severity) that will result from each failure mode that may occur.
- Calculate the criticality for each potential failure mode by obtaining the product of the three factors:
Mode Criticality = Item Unreliability x Mode Ratio of Unreliability x Probability of Loss
- Calculate the criticality for each item by obtaining the sum of the criticalities for each failure mode that has been identified for the item.
Item Criticality = SUM of Mode Criticalities
Qualitative Criticality Analysis Method.
- Rate the severity of the potential effects of failure.
Rate the likelihood of occurrence for each potential failure mode.
Compare failure modes via a Criticality Matrix, which identifies severity on the horizontal axis and occurrence on the vertical axis.
Software FMECA
- 1. Break the software into logical components, such as functions or tasks
2. Determine potential failure modes for each component
- (e.g., “software fails to respond” or “software responds with wrong value”)
3. Using a failure mode table, use the failure modes to fill in the Software FMECA worksheet
4. Determine all possible causes for each failure mode
- (e.g. “logic omitted/implemented incorrectly”, “incorrect coding of algorithm”, “incorrect manipulation of data”, etc),
5. Determine the effects of each failure mode at progressively higher levels (assume inputs to the software are good)
6. Assign failure rates; occurrence, severity and detectability values; and calculate the Risk Priority Number (RPN)
7. Identify all corrective actions (CAs) that should be, or have been, implemented, plus all open item
sFMECA Table
- Index No
- Unit
- Function
- Failure Mode
- Possible Failure Causes
- Effects on
- Unit
- SubSystem
- System
- Fail Rate
- Occurence
- Severity
- RPN
- 1. is an extension of the FMEA.
- FTA: Fault Tree Analysis
- 1. systematic methodology for defining a specific undesirable event (normally a mission failure) and determining all the possible reasons that could cause the event to occur
2. logical, top-down method aimed at analyzing the effects of initiating faults and events upon a complex system
3. utilizes a block diagram approach that displays the state of the system in terms of the states of its components.
4. tasks are typically assigned to engineers having a good understanding of the interoperability of the system and subsystem level components
5. The top level events in the Fault Tree are typically mission essential functions (MEFs) of the system, which should be defined in the system‟s Failure Definition / Scoring Criteria (FD/SC).
6. If a component failure results in a loss of one of these mission essential functions, then that failure would be classified as an Operation Mission Failure (OMF), System Abort (SA) failure,
7. The idea behind the Fault Tree is to break down each one of the MEFs further and further until all failure causes of components and subcomponents are captured.
FTA Software
- Fault Tree Analysis for Systems Containing Software.
- To be most effective, a system-level FTA must consider software, hardware, interfaces and human interactions.
Top-level undesired events may be derived from:
- 1. Engineering judgment based on what the system should ultimately not do
- 2. Problems known from previous experience, or from historical “clue lists”
- 3. Outputs of a FMEA – or a FMECA
- 4. Information from a Preliminary Hazard Analysis
- To be most effective, a system-level FTA must consider software, hardware, interfaces and human interactions.
- The basic steps for performing a FTA for software (same as for hardware) are as follows:
- 1. Definition of the problem
- 2. Analysis Approach
- a. Identify the top-level event
- b. Develop the fault tree
- c. Analyze the fault tree
- i. Delineate the minimum cut set
- ii. Determine the reliability of the top-level event
- iii. Review analysis output
- Fault Tree Analysis for Systems Containing Software.
- 1. systematic methodology for defining a specific undesirable event (normally a mission failure) and determining all the possible reasons that could cause the event to occur
- FRB: Failure Review Board
- FPRB: Failure Prevention Review Board
- 1. The FPRB is chaired by the Program Manager and generally incorporates key members from across several areas of the program development process.
2. The board reviews and approves each proposed corrective action and determines the effectiveness of the fix by assigning a fix effectiveness factor.
3. If the corrective action is approved, the fixes are implemented and effectiveness verified.
4. Should the corrective actions not be approved, the root causes are analyzed to determine improvements or other potential fixes
- 1. The FPRB is chaired by the Program Manager and generally incorporates key members from across several areas of the program development process.
- FRACAS: Failure Reporting and Corrective Action System
- Purpose: closed-loop process whose purpose is to provide a systematic way to
- 1. report
- 2. organize
- 3. analyze failure data
- Failure Reporting.
- – Failures and faults that occur during developmental or operational testing or during inspections are reported.
– The failure report should include
- 1. identification of the failed item,
- 2. symptoms of the failure
- 3. testing conditions,
- 4. item operating conditions
- 5. time of the failure.
- – Failures and faults that occur during developmental or operational testing or during inspections are reported.
- Failure Analysis.
- – Each reported failure is analyzed to determine the root cause.
– A FRACAS should have the capability to document the results and conclusions of the root cause analysis.
- – Each reported failure is analyzed to determine the root cause.
- Design Modifications.
- – Details of the implemented corrective actions for each failure mode should be documented. The precise nature of the design change should be captured as well as date of implementation.
- Failure Verification.
- – All reported failures need to be verified as actual failures by repeating the failure or evidence of failure, such as a leak or damaged hardware.
– This verification needs to be tracked within the FRACAS.
- – All reported failures need to be verified as actual failures by repeating the failure or evidence of failure, such as a leak or damaged hardware.
- Purpose: closed-loop process whose purpose is to provide a systematic way to
- FTA Example:
- Wafer Unable To Move To The Next Step
- Loss of Wafer Location
- Wafer Sensor Detection Failed
- Wafer Sensor Board Failed
- Loss of Wafer: Wafer Break
- Actual Wafer Break
- Partial Wafer Break
- Loss of Wafer Location
- Wafer Unable To Move To The Next Step
- MEF: Mission Essential Function
- OMF: Operation Mission Failure
- SA: System Abort
- FD: Failure Definition
- SC: Scoring Criteria
- Book: Design For Reliability
- About the Book
- Title: Design For Relibility
- Authors: Raheja/Gullo
- Chapters: 18
- Pages: 334
- 1 Design for Reliability Paradigms 1
- Dev Raheja
- Why Design for Reliability? 1
- 1. The science of reliability has not kept pace with user expectations.
2. Many corporations still use MTBF (mean time between failures) as a measure of reliability, which, depending on the statistical distribution of failure data, implies acceptance of roughly 50 to 70% failures during the time indicated by the MTBF.
3. No user today can tolerate such a high number of failures. Ideally, a user does not want any failures for the entire expected life!
4. The life expected is determined by the life inferred by users, such as 100,000 miles or 10 years for an automobile, at least 10 years for kitchen appliances, and at least 20 years for a commercial airliner.
5. Most commercial companies, such as automotive and medical device manufacturers, have stopped using the MTBF measure and aim at 1 to 10% failures during a self-defined time. This is still not in line with users’ dreams.
6. The real question is: Why not design for zero failures if we can increase profits and gain more market share? Zero failures implies zero mission-critical failures or zero safety-critical system failures.
7. As a minimum, systems in which failures can lead to catastrophic consequences must be designed for zero failures.
8. There are companies that are able to do this. Toyota, Apple, Gillette, Honda, Boeing, Johnson & Johnson, Corning, and Hewlett-Packard are a few examples.
9. The aim of design for reliability (DFR) is to design-out failures of critical system functions in a system. The number of such failures should be zero for the expected life of the product.
10. Some components may be allowed to fail, such as in redundant systems
- 1. The science of reliability has not kept pace with user expectations.
- Reflections on the Current State of the Art 2
- 1. Reliability is defined as the probability of performing all the functions (including safety functions) satisfactorily for a specified time and specified use conditions.
2. The functions and use conditions come from the specification.
3. If a specification misses or is vague 60% or more of the time, the reliability predictions are of very little value. This is usually the case
4.The second big issue is: How many failures should be tolerable? Some readers may not agree that we can design for zero critical failures, but the evidence supports the contrary conclusion. We may not be able to prevent failures that we did not foresee, but we can design out all the critical failure modes that we discover during the requirements analysis and in the failure mode and effects analysis (FMEA).
5. In over 30 years’ experience, I have yet to encounter a failure mode that cannot be designed-out. The cost is usually not an issue if the FMEA is conducted and the improvements are made during the early design stage.
6. The time specified for critical failures in the reliability definition should be the entire lifetime expected.
- 1. Reliability is defined as the probability of performing all the functions (including safety functions) satisfactorily for a specified time and specified use conditions.
- The Paradigms for Design for Reliability 4
- Summary 13
- References 13
- 2 Reliability Design Tools 15
- Joseph A. Childs
- Introduction 15
- Reliability Tools 19
- Test Data Analysis 31
- Summary 34
- References 35
- 3 Developing Reliable Software 37
- Samuel Keene
- Introduction and Background 37
- Software Reliability: Definitions and Basic Concepts 40
- Software Reliability Design Considerations 44
- Operational Reliability Requires Effective Change Management 48
- Execution-Time Software Reliability Models 48
- Software Reliability Prediction Tools Prior to Testing 49
- References 51
- 4 Reliability Models 53
- Louis J. Gullo
- Introduction 53
- Reliability Block Diagram: System Modeling 56
- Example of System Reliability Models Using RBDs 57
- Reliability Growth Model 60
- Similarity Analysis and Categories of a Physical Model 60
- Monte Carlo Models 62
- Markov Models 62
- References 64
- 5 Design Failure Modes, Effects, and Criticality Analysis 67
- Louis J. Gullo
- Introduction to FMEA and FMECA 67
- Design FMECA 68
- Principles of FMECA-MA 71
- Design FMECA Approaches 72
- Example of a Design FMECA Process 74
- Risk Priority Number 82
- Final Thoughts 86
- References 86
- 6 Process Failure Modes, Effects, and Criticality Analysis 87
- Joseph A. Childs
- Introduction 87
- Principles of P-FMECA 87
- Use of P-FMECA 88
- What Is Required Before Starting 90
- Performing P-FMECA Step by Step 91
- Improvement Actions 98
- Reporting Results 100
- Suggestions for Additional Reading 101
- 7 FMECA Applied to Software Development 103
- Robert W. Stoddard
- Introduction 103
- Scoping an FMECA for Software Development 104
- FMECA Steps for Software Development 106
- Important Notes on Roles and Responsibilities with Software
- FMECA 116
- Lessons Learned from Conducting Software FMECA 117
- Conclusions 119
- References 120
- 8 Six Sigma Approach to Requirements Development 121
- Samuel Keene
- Early Experiences with Design of Experiments 121
- Six Sigma Foundations 124
- The Six Sigma Three-Pronged Initiative 126
- The RASCI Tool 128
- Design for Six Sigma 129
- Requirements Development: The Principal Challenge to System
- Reliability 130
- The GQM Tool 131
- The Mind Mapping Tool 132
- References 135
- 9 Human Factors in Reliable Design 137
- Jack Dixon
- Human Factors Engineering 137
- A Design Engineer’s Interest in Human Factors 138
- Human-Centered Design 138
- Human Factors Analysis Process 144
- Human Factors and Risk 150
- Human Error 150
- Design for Error Tolerance 153
- Checklists 154
- Testing to Validate Human Factors in Design 154
- References 154
- 10 Stress Analysis During Design to Eliminate Failures 157
- Louis J. Gullo
- Principles of Stress Analysis 157
- Mechanical Stress Analysis or Durability Analysis 158
- Finite Element Analysis 158
- Probabilistic vs. Deterministic Methods and Failures 159
- How Stress Analysis Aids Design for Reliability 159
- Derating and Stress Analysis 160
- Stress vs. Strength Curves 161
- Software Stress Analysis and Testing 166
- Structural Reinforcement to Improve Structural
- Integrity 167
- References 167
- 11 Highly Accelerated Life Testing 169
- Louis J. Gullo
- Introduction 169
- Time Compression 173
- Test Coverage 174
- Environmental Stresses of HALT 175
- Sensitivity to Stresses 176
- Design Margin 178
- Sample Size 180
- Conclusions 180
- Reference 181
- 12 Design for Extreme Environments 183
- Steven S. Austin
- Overview 183
- Designing for Extreme Environments 183
- Designing for Cold 184
- Designing for Heat 186
- References 191
- 13 Design for Trustworthiness 193
- Lawrence Bernstein and C. M. Yuhas
- Introduction 193
- Modules and Components 196
- Politics of Reuse 200
- Design Principles 201
- Design Constraints That Make Systems Trustworthy 204
- Conclusions 210
- eferences and Notes 211
- 14 Prognostics and Health Management Capabilities to Improve Reliability 213
- Louis J. Gullo
- Introduction 213
- PHM Is Department of Defense Policy 216
- Condition-Based Maintenance vs. Time-Based Maintenance 216
- Monitoring and Reasoning of Failure Precursors 217
- Monitoring Environmental and Usage Loads for Damage
- Modeling 218
- Fault Detection, Fault Isolation, and Prognostics 218
- Sensors for Automatic Stress Monitoring 220
- References 221
- 15 Reliability Management 223
- Joseph A. Childs
- Introduction 223
- Planning, Execution, and Documentation 229
- Closing the Feedback Loop: Reliability Assessment, Problem Solving, and Growth 232
- References 233
- 16 Risk Management, Exception Handling and Change Management 235
- Jack Dixon
- Introduction to Risk 235
- Importance of Risk Management 236
- Why Many Risks Are Overlooked 237
- Program Risk 239
- Design Risk 241
- Risk Assessment 242
- Risk Identification 243
- Risk Estimation 244
- Risk Evaluation 245
- Risk Mitigation 247
- Risk Communication 248
- Risk and Competitiveness 249
- Risk Management in the Change Process 249
- Configuration Management 249
- References 251
- 17 Integrating Design for Reliability with Design for Safety 253
- Brian Moriarty
- Introduction 253
- Start of Safety Design 254
- Reliability in System Safety Design 255
- Safety Analysis Techniques 255
- Establishing Safety Assessment Using the Risk Assessment Code
- Matrix 260
- Design and Development Process for Detailed Safety Design 261
- Verification of Design for Safety Includes Reliability 261
- Examples of Design for Safety with Reliability Data 262
- Final Thoughts 266
- References 266
- 18 Organizational Reliability Capability Assessment 267
- Louis J. Gullo
- Introduction 267
- The Benefits of IEEE 1624-2008 269
- Organizational Reliability Capability 270
- Reliability Capability Assessment 271
- Design Capability and Performability 271
- IEEE 1624 Scoring Guidelines 276
- SEI CMMI Scoring Guidelines 277
- Organizational Reliability Capability Assessment Process 278
- Advantages of High Reliability 282
- Conclusions 283
- References 284
- About the Book
- Compiled By:
- Rigel Arcayan
- blindcaveman.wordpress.com