Aviation Training Experts™

handbook

Aviation Instructor’s Handbook

FAA-H-8083-9B Version 2020

Chapter 6

Assessment

Figure 6-3. Effective tests have six primary characteristics.
Figure 6-3. Effective tests have six primary characteristics.

Reliability is the degree to which test results are consistent with repeated measurements. If identical measurements are obtained every time a certain instrument is applied to a certain dimension, the instrument is considered reliable. The reliability of a written test is judged by whether it gives consistent measurement to a particular individual or group. Keep in mind, though, that knowledge, skills, and understanding can improve with subsequent attempts at taking the same test, because the first test serves as a learning device.

Validity is the extent to which a test measures what it is supposed to measure, and it is the most important consideration in test evaluation. The instructor must carefully consider whether the test actually measures what it is supposed to measure. To estimate validity, several instructors read the test critically and consider its content relative to the stated objectives of the instruction. Items that do not pertain directly to the objectives of the course should be modified or eliminated.

Usability refers to the functionality of tests. A usable written test is easy to give if it is printed in a type size large enough for learners to read easily. The wording of both the directions for taking the test and of the test items needs to be clear and concise. Graphics, charts, and illustrations appropriate to the test items must be clearly drawn, and the test should be easily graded.

Objectivity describes singleness of scoring of a test. Essay questions provide an example of this principle. It is nearly impossible to prevent an instructor’s own knowledge and experience in the subject area, writing style, or grammar from affecting the grade awarded. Selection-type test items, such as true/false or multiple choice, are much easier to grade objectively.

Comprehensiveness is the degree to which a test measures the overall objectives. Suppose, for example, an AMT wants to measure the compression of an aircraft engine. Measuring compression on a single cylinder would not provide an indication of the entire engine. Similarly, a written test must sample an appropriate cross-section of the objectives of instruction. The instructor makes certain the evaluation includes a representative and comprehensive sampling of the objectives of the course.

Discrimination is the degree to which a test distinguishes the difference between learners and may be appropriate for assessment of academic achievement. However, minimum standards are far more important in assessments leading to pilot certification. If necessary for classroom evaluation of academic achievement, a test must measure small differences in achievement in relation to the objectives of the course. A test designed for discrimination contains:

  1. A wide range of scores
  2. All levels of difficulty
  3. Items that distinguish between learners with differing levels of achievement of the course objectives

Please see Appendix B for information on the advantages and disadvantages of multiple choice, supply type, and other written assessment instruments, as well as guidance on creating effective test items.

Authentic Assessment

Authentic assessment asks the learner to perform real-world tasks and demonstrate a meaningful application of skills and competencies. Authentic assessment lies at the heart of training today’s aviation learner to use critical thinking skills. Rather than selecting from predetermined responses, learners must generate responses from skills and concepts they have learned. By using open-ended questions and established performance criteria, authentic assessment focuses on the learning process, enhances the development of real-world skills, encourages higher order thinking skills, and teaches learners to assess their own work and performance.

Learner-Centered Assessment

There are several aspects of effective authentic assessment. The first is the use of open-ended questions in what might be called a “collaborative critique,” which is a form of learner-centered grading. As described in the scenario that introduced this chapter, the instructor begins by using a four-step series of open-ended questions to guide the learner through a complete self-assessment.

Replay—the instructor asks the learner to verbally replay the flight or procedure. While the learner speaks, the instructor listens for areas where the account does not seem accurate. At the right moment, the instructor discusses any discrepancy with the learner. This approach gives the learner a chance to validate his or her own perceptions, and it gives the instructor critical insight into the learner's judgment abilities.

Reconstruct—the reconstruction stage encourages learning by identifying the key things that the learner would have, could have, or should have done differently during the flight or procedure.

Reflect—insights come from investing perceptions and experiences with meaning, requiring reflection on the events. For example:

  1. What was the most important thing you learned today?
  2. What part of the session was easiest for you? What part was hardest?
  3. Did anything make you uncomfortable? If so, when did it occur?
  4. How would you assess your performance and your decisions?
  5. How did your performance compare to the standards in the ACS?

Redirect—the final step is to help the learner relate lessons learned in this session to other experiences and consider how they might help in future sessions. Questions might include:

  • How does this experience relate to previous lessons?
  • What might be done to mitigate a similar risk in a future situation?
  • Which aspects of this experience might apply to future situations, and how?
  • What personal minimums should be established, and what additional proficiency flying and/or training might be useful?

Any self-assessment stimulates growth in the learner’s thought processes and, in turn, behaviors. An in-depth discussion between the instructor and the learner may follow, which compares the instructor’s assessment to the learner’s self-assessment. Through this discussion, the instructor and the learner jointly determine the learner’s progress. The progress may be recorded on a rubric as part of a training program. As explained earlier, a rubric is a guide for scoring performance assessments in a reliable, fair, and valid manner. It is generally composed of dimensions for judging learner performance, a scale for rating performances on each dimension, and standards of excellence for specified performance levels.