Skip to main content

TECHNIQUES OF EVALUATION

 TECHNIQUES OF EVALUATION

OBSERVATION:

Introduction and Meaning of Observation
          The most common method used for getting information about the various things around us, is to observe those things and also the various processes related to those things. Hence, it can be said that observation acts as a fundamental and the basic method of getting information about anything. But it must be kept in mind that observation is not just seeing things but it is carefully watching the things and trying to understand them in depth, in order to get some information about them.

Nature of Observation technique

1.       Systematic and Purposeful: It is a planned, deliberate process aimed at specific learning objectives, distinguishing it from casual watching.

2.       Real-time Evaluation: It captures student behaviour, skills, and reactions instantly as they naturally occur in the classroom, lab, or playground.

3.       Measures the affective domain: It is highly effective for evaluating non-cognitive traits like interest, motivation, values, and scientific attitudes.

4.       Assesses Psychomotor skills: It directly measures hands-on practical abilities, such as handling laboratory apparatus, drawing diagrams or performing physical tasks.

5.       Natural Environment: It takes place in the student’s everyday learning environment, reducing the artificial stress often caused by formal written tests.

6.       Continuous Process: It allows for ongoing, formative assessment over time rather than providing just a single “Snapshot” grade.

7.       Bridges Qualitative and Quantitative Data: While it begins as qualitative descriptions of behavior, tools like rating scales can convert observations into quantifiable data.

8.       Relies on Structured Tools: It depends on specific instruments-such as checklists, rating scales, and anecdotal records to maintain focus and consistency.

9.       Prone to subjectivity: The greatest limitation is personal bias, as different observers might interpret the exact same student behavior differently.

10.   The Hawthorne Effect: Its accuracy can be compromised if students notice they are being watched, causing them to alter their natural behavior.

Construction of Observation schedule

1.       Clear Objectives: Start with a well-defined purpose, specifying exactly what skill, concept, or behavior needs evaluation.

2.       Measurable Behaviors: Break down broad goals into specific, observable, and concrete actions (e.g., “pours liquid carefully” rather than “is good at experiments”).

3.       Appropriate Format selection: Choose a structured recording method-such as a checklist for presence/absence or a rating scale for quality-that fits the objective.

4.       Logical Sequencing: Arrange behavioral items chronologically in the exact order they are naturally expected to occur during the activity.

5.       Simplicity and Brevity: Keep the layout clean and limit the number of items so the observer can record details quickly without missing teal-time actions.

6.       Unambiguous language: Use clear, precise terminology for each item to prevent confusion or varied interpretations by different observers.

7.       Comprehensive Metadata: Include dedicated spaces for essential contextual details, such as data, time subject, class, and student identifiers.

8.       Explicit scoring Guides: Provide brief, clear instructions on how to use the scale or checkmarks to ensure consistency throughout the evaluation.

9.       Pilot testing: Try out the draft schedule in a trial session to spot flaws, timing issues or vague items before using it formally.

10.   Refinement: Revise and finalize the tool based on trial feedback to minimize observer bias and maximize the tool’s reliability.

Uses of Observation Technique

1.       Assessing Practical Skills: It directly evaluates hands-on performance and psychomotor skills, such as handling apparatus in a science lab, using equipment or drawing diagrams.

2.       Evaluating Affective Attributes: It is the primary tool for measuring non-cognitive traits that written tests miss, such as student interests, values, teamwork and scientific attitudes.

3.       Continuous formative assessment: It allows teachers to monitor ongoing learning progress daily in a natural classroom environment, providing immediate feedback to guide instruction. 

4.       Diagnostic tool for learning difficulties: It helps identify specific operational errors, behavioral issues, or conceptual hurdles a student faces during real-time tasks.

5.       Evaluating young children: It serves as an indispensable evaluation tool in early childhood education where students cannot yet read or write formal examinations.

6.       Assessing Group Dynamics and social skills: It captures how students interact, communicate, share responsibilities, and exhibit leadership during collaborative or co-curricular activities.

7.       Curriculum and Instructional feedback: It reveal how well students respond to a specific teaching method, multimedia tool, or lesson plan, allowing the teacher to adjust their approach.

8.       Validating Self-report Data: It acts as a reliable cross-reference to verify if a student’s actual classroom behavior aligns with what they claim on questionnaires or self-assessments.

9.       Studying Natural Behavioral pattern: It records authentic, unprompted student behavior in unstructured settings like the playground, library or school exhibitions.

10.   Research and Action Research: It provides qualitative, real-time data for classroom teachers conducting action research to solve immediate educational problems and improve classroom management.

Advantages of Observation
1. Very direct method for collecting data or information – best for the study of human behavior.
2. Data collected is very accurate in nature and also very reliable.
3. Improves precision of the research results.
4. Problem of depending on respondents is decreased.
5. Helps in understanding the verbal response more efficiently.
6. By using good and modern gadgets – observations can be made continuously and also for a larger duration of time period.
7. Observation is less demanding in nature, which makes it less bias in working abilities.
8. By observation, one can identify a problem by making an in-depth analysis of the problems.

Disadvantages of Observation 
1. Problems of the past cannot be studied by means of observation.
2. Having no other option, one has to depend on the documents available.
3. Observations like the controlled observations require some especial instruments or tools for effective working, which are very much costly.
4. One cannot study opinions by this means.
5. Attitudes cannot be studied with the help of observations.
6. Sampling cannot be brought into use.
7. Observation involves a lot of time as one has to wait for an event to happen to study that particular event.
8. The actual presence of the observer himself Vis a Vis the event to occur is almost unknown, which acts as a major disadvantage of observation.
9. Complete answer to any problem or any issue cannot be obtained by observation alone.

 

Suggestions to help make valid observations

1.       Plan in advance what is to be observed.

2.       The observer must be cognizant of sampling errors. There should be frequent, short observation distributed over a period of several weeks and at different times of the day.

3.       Co-ordinate the observations with your teaching. Otherwise, there is great danger that invalid observations will result.

4.       Record and summarize the observation immediately after it has occurred. More important, however, is the fact that when pupils know they are being observed, their resultant behaviour maybe atypical.

5.       Make no interpretations concerning the behaviour until later on. Otherwise, it may interfere with the objectivity of gathering observational data.

6.       Prepare some sort of list, guide or form to help make the observation process objective and systematic.

QUESTIONNAIRE

        A questionnaire is a list of planned written questions related to a particular topic or series of topics. Space is provided for the reply to each question.

          In Structured (closed-end) type of questionnaire, the answers are checked or underlined by the respondent. In the unstructured (open-end) type, the respondent is allowed to make free responses to the questions. The inventory comes under the first type.

Nature of questionnaire

1.       It is a formal research instrument consisting o predefined set of questions designed to gather data from respondents systematically.

2.       Every participant is asked the exact same questions in the exact same order, ensuring consistency across all responses.

3.       It eliminates interviewer bias since the researcher isn’t actively guiding the conversation, allowing respondents to answer more freely. 

4.       Every single question is strictly tied to a specific research objective; it leaves no room for irrelevant or filler questions.

5.       It can gather statistical data using closed-ended questions (like multiple-choice) or deep insights using open-ended questions.

6.       It can be distributed to hundreds or thousands of people simultaneously, especially through online tools, making it highly cost-effective.

7.       Because respondents can often complete them anonymously, they are more likely to share honest answers on sensitive topics.

8.       It follows a psychological sequence, usually starting with easy, engaging questions to build comfort, before moving to complex or personal ones.

9.       It is not final until it is pre-tested on a small group to catch confusing wording, errors, or technical bugs.

10.   Its success relies heavily on the willingness o respondents to complete it; therefore, clarity and brevity are crucial.

Construction of Questionnaire tool

1.       Define clear assessment objectives: Before writing a single question, map out exactly what knowledge, skills, or competencies you are testing. Every item must align directly with your predefined learning outcome.

2.       Map items to a table of specifications (blue-print): Create a matrix that crosses your learning objectives with the cognitive levels being tested (e.g., recall, application, analysis.) This ensures balanced coverage of the curriculum and prevents over-testing one specific topic.

3.       Choose the right question formats: Select item types that best fit the objective. Use objective formats (multiple choice, matching, true/false) for broad coverage and efficient grading and subjective formats (short answer, Essay) to evaluate higher order thinking and synthesis.  

4.       Ensure clarity and Conciseness: Write stems (the question part) and prompts using simple, unambiguous language. Avoid double negatives, complex sentence structures, or unnecessary fluff that tests a student’s reading comprehension rather than their actual knowledge.

5.       Follow strict Multiple-choice guidelines: Ensure there is only one demonstrably correct answer. Distractors (wrong options) should be plausible, homogeneous in length and grammar and free of giveaway clues like “all of the above “ or never/always”.

6.       Sequence Questions strategically: Organize the questionnaire logically. Group similar question formats together, and arrange items from easiest to hardest. Starting with accessible questions helps reduce test anxiety and builds student confidence.

7.       Establish standardized Instructions: Provide clear, explicit directions for each section. Specify how students should record their answers, the time limit, and the grading criteria (e/g., whether guessing is penalized or how partial credit is awarded).

8.       Create Detailed scoring rubric or answer key: Develop your grading criteria simultaneously with the questions. For essays or short answers, a structured rubric ensures consistency and fairness during grading. For objective items, ensure your answer key is error free.  

9.       Review and Pre-test (pilot) the tool: Have a peer or subject-matter expert review the questionnaire to catch typos, ambiguities, or formatting errors, if possible, run a small pilot test to ensure the timing is appropriate and instructions are clear.

10.   Plan for Post-assessment analysis: After administration, analyze the tool’s performance using metrics like item difficulty and discrimination. This helps identify flawed questions that should be dropped or modified for future assessments.

Uses of questionnaire tool

1.       Evaluating Student Feedback: Questionnaires are widely used to gather student feedback regarding teaching methods, course content, and classroom environment, helping educators refine their instructional strategies.

2.       Formative assessment and diagnostic screening: they help assess students’ prior knowledge, interests, and learning styles before a unit begins, or gauge their ongoing understanding during a course to identify learning gaps.

3.       Measuring Attitudes and motivation: Beyond academic skills, questionnaires effectively assess non-cognitive domains, such as student’s attitude toward a subject (e.g., science or maths) their motivation levels, and academic anxiety.

4.       Large-scale Data collection: they allow researchers and administrator to gather assessment data from hundreds of participants simultaneously, making them highly cost-effective and time-efficient compared to face-to-face interviews.

5.       Action research support: For classroom teachers conducting action research, questionnaires serve as a primary data collection tool to evaluate the impact of a new teaching intervention or multimedia approach.

6.       Standardized and objective scoring: Closed-ended questionnaires (like those using Likert scales) provide quantitative data that can be objectively scored, aggregated, and statistically analyzed with minimal grader bias.

7.       Self-assessment and metacognition: When designed for students to reflect on their own learning habits, questionnaires encourage metacognition-helping learners evaluate their own study strategies, strengths, and weaknesses.

8.       Peer Evaluation: In collaborative or project-based learning, structured questionnaires enable students to assess their peers’ contributions, teamwork skills, and effort anonymously and objectively.

9.       Institutional and Program Evaluation: Educational institutions use questionnaires to conduct broader curriculum evaluations, stakeholder satisfaction surveys (from alumni or employers) and accreditation reviews.

10.   Parent-teacher Insights: They bridge the communication gap by gathering insights from parents regarding a child’s behavior at home, study habits, and emotional well-being, providing a holistic view of the learner.

CHECKLIST

         A check list consists of a listing of steps, activities or behaviour which the observer records when an incident occurs. It is similar in appearance and use to a rating scale and is classified by some as a type of rating scale.

           A check list enables the observer to note only whether or not a trait or characteristic is present. It does not permit the observer to rate the quality of a particular behaviour or its frequency of occurrence or the extent to which a particular characteristic is present. When such information is desired, the check list is definitely inappropriate.

Nature of Checklist

1.       Binary/dichotomous: It operates on a two-choice system (yes/No, Present/Absent) with no middle ground.

2.       Objective and observable: It focus purely on concrete, visible behaviors or items, significantly reducing evaluator bias.

3.       Fixed and Structured: It consists of a predetermined list of specific criteria that remains identical for every individual being assessed.

4.       Performance tracking: It captures whether a specific step in a process was followed or if a specific feature in a product exists.

5.       Non-qualitative: It records the presence of an action or item, not the quality or proficiency of how well it was executed.

CONSTRUCTION OF CHECKLIST TOOL

1.       Define the purpose: Clearly identify the specific skill, behavior or product you want to assess (e.g., formulating a hypothesis in a science lab).

2.       Break down into steps: Split the task into distinct, concrete, and observable actions. Avoid vague ideas; focus on what you can see.

3.       Keep items Brief and positive: Write short, clear statements stating what the learner should do, rather than what they shouldn’t (e.g., “Wears safety goggles” instead of “Does not forget goggles”).

4.       Sequence logically: Arrange the items in the exact chronological order in which they should naturally occur during the performance.

5.       Add a binary scale: Provide a simple, two-choice response column next to each item, such as yes/No, present/absent or Achieved/Not achieved.

 

A Specimen of Check list

Directions:  Listed below are a series of characteristics related to health practices. Check those characteristics which are applicable to students.

Characteristics to be observed                                        Roll Nos. of the pupils.

                                                                                         1       2          3             4          5            6        7  

1.       Take a balanced diet                   

2.       Washes hands before breakfast

3.       Brushes teeth after eating

4.       Drinks plenty of water at the time of eating

5.       Brushes teeth before going to bed

6.       Goes for a walk daily; etc

Uses/Advantages of Check lists

1.       They are adaptable to most subject-matter areas.

2.       They are useful in evaluating those learning activities that involve a product, process and some aspects of personal-social adjustment.

3.       They are most useful for evaluating those processes that can be sub-divided into a series of clear, distinct, separate actions.

4.       When properly prepared, they constrain the observer to direct his attention to clearly specified traits or characteristics.

5.       They allow inter-individual comparisons to be made on a common set of traits or characteristics.

6.       They provide a simple method to record observations.

7.       They objectively evaluate traits or characteristics.

RATING SCALE

           Rating scales resemble check lists but are used when finer discriminations are required. Instead of merely indicating the presence or absence of a trait or characteristic, it enables us to indicate the degree to which a trait is present. Rating scales provide systematic procedures for obtaining, recording and reporting the observer’s judgements. That may be filled out while the observation is made, immediately after the observation is made or, as often is the case long after the observation.

Nature of Rating scale

1.       Measures Intensity, not just categories: Unlike “yes/No” or multiple-choice questions that categorize data, a rating scale measures the degree, frequency or strength of an attribute (e.g., how much a student understands a concept, or how often a behavior occurs).

2.        Built on a defined continuum: they rely on a structured, sequential progression of options. Responses move logically from one extreme to the other (e.g., from strongly Disagree to strongly Agree, or never to always).

3.       Yields quantifiable Easy-to-analyze data: By assigning numbers to subjective human options or observations (like coding responses from 1 to 5), they turn qualitative feelings into quantitative data that can be easily graphed, averaged and statistically analyzed.

4.       Subject to structural choices (Even vs Odd): The design of the scale dictates how respondents answer. Including an odd number of points provides a neutral midpoint, while using an even number forces a choice, preventing respondents from sitting on the fence.

5.       Highly vulnerable to human bias: Because they rely on human judgment, they are constantly prone to errors like central tendency (everyone playing it safe and picking the idle numbers) or social desirability (choosing the answer that looks best rather than the truth).

Types of Rating Scales:

a)      Numerical Rating Scale:

     This is one of the simplest types of rating scales. The rater simply marks a number that indicates the extent to which a characteristic or trait is present. The trait is presented as a statement and values from 1 to 5 (a maximum of 10) are assigned to each trait that is rated. Typically a common key is used throughout, the key providing a verbal description.

Direction: Encircle the appropriate number showing the extent to which the pupil exhibits his skill in questioning.

Key: 5-outstanding, 4-above average, 3-average, 2-below average, 1- unsatisfactory.

Skill:

1.       Questions were specific:                                     1       2       3         4        5

2.       Questions were relevant to the

                            Topic discussed.                        1       2       3         4        5

3.       Questions were grammatically correct, etc.    1       2       3         4        5

b)      Graphic Rating Scale:

As in the case of the numerical rating scale, the rater is required to assign some value to a specific trait. This time, however, instead of using predetermined scale values, the ratings are made in a graphic form-a position anywhere along a continuum.

Direction: Rate for each characteristic listed below along the continuum from 1 to 5. You can use points between the scale values. Mark X at the appropriate place along the continuum.

1.       Were the illustrations used interesting?

     1                                2                            3                                      4                       5

Too little                   Little                   Adequate                         Much              Too much

2.       How attentive were you in the class?

     1                              2                                3                         4                             5

                            Very inattentive          Inattentive                                         Attentive                  Very attentive

3.       Did the speech show good organisaqtion?

_____________________________________________________________________

    1                                2                              3                              4                               5

Very poor                                               Average                                                  Very good

Advantage: If a number of traits are rated on the same page with a common set of categories, a behavioural profile can be constructed.

C) Descriptive Graphic Rating Scale

          This type of scale is generally the most desirable type of scale to use.

Directions: As shown above for the graphic rating scale.

1.       While preparing a blackboard summary, how was the penmanship?

 

Legible, beautiful,                 normally readable,                                         illegible, bad-looking

Uniform size and                   good-looking,                                                   tends to draw outlines

Slant                                       fluent motion

    Such specific descriptions contribute to a greater objectivity of the rating process. The description also helps to clarify and further define a particular dimension.

d) Ranking:

In the ranking procedure, the rater, instead of assigning a numerical value to each student with regard to a characteristic, ranks a given set of individuals from high to low on the characteristic this rated. To ensure that the pupils are validly ranked, rank from the both extreme towards the middle. This simplifies the task of the teacher. The ranking procedure becomes very cumbersome when a large number of students or characteristics per student are to be ranked.

 

Construction of Rating scale

1.       Define the Trait: Clearly identify and isolate the specific concept attitude, or behavior you intend to measure so the scale remains strictly focused.

2.       Choose the Format: Decide on the number of scale points and whether to use an odd number to allow neutrality or an even number to force a choice.

3.       Draft the items: Write simple, single-focused statements that avoid double-barreled ideas and include a mix of positive and negative wording.

4.       Anchoring the points: Attach precise, logical descriptions to each numerical value so every respondent interprets the progression identically.

5.       Pre-test the scale: Pilot the scale on a small test group to catch confusing language and eliminate formatting flows before final deployment.

Uses of Rating scale as tool

1.       They measure specified outcomes or objectives of education deemed to be significant or important to the teacher.

2.       They evaluate procedures (such as paying on an instrument, working in the laboratory, typing, cooking, singing, oral reading, acting in a play), Products (such as typed letters, a speech, written themes, samples of handwriting, art work), and personal social development.

3.       They help teachers to rate their students periodically on various characteristics such as punctuality, enthusiasm, cheerfulness, co-cooperativeness, consideration for others and other personality traits.

4.       They can also be used by pupil to rate himself.

INTERVIEW

         The word interview comes from Latin and middle French words meaning to “see between” or “see each other”. Generally, interview means a private meeting between people when questions are asked and answered. The person who answers the questions of an interview is called in interviewer. The person who asks the questions of our interview is called an interviewer.

Definitions

`1.  According to Gary Dessler, “An interview is a procedure designed to obtain information from a person’s oral response to oral inquiries.”

2. According to Thill and Bovee, “An interview is any planed conversation with a specific purpose involving two or more people”.

Nature of Interview

1.       Interactive and Flexible: It is a dynamic, two-way conversation that allows the assessor to probe deeper, clarify questions, and adapt to the candidate’s responses in real time.

2.       Holistic Evaluation: It measures complex skills that paper tests miss, such as verbal communication, interpersonal skills, confidence and critical thinking under pressure.

3.       Versatile purpose: It can be used formatively (to diagnose learning gaps and thought processes) or combatively (like a viva voce or a hiring panel to make a final decision).

4.       Varying structure: Its reliability depends on depends on design:

·         Structured: Highly objective and easy to compare, but rigid.

·         Semi-structured: Balance, allowing flexible follow-up while staying on topic.

·         Unstructured: Rich in qualitative detail, but highly subjective and difficult to score.

5.       Prone to Bias: Because human judgment is involved, it requires standardized rubrics or multiple interviewers to stay fair and valid.

Construction of Interview

1.       Define the objectives: Clearly identify what you want to measure (e.g., specific knowledge, communication skills, or critical thinking).

2.       Choose the Structure: Decide whether a structured approach (fixed questions for exact comparison) or a semi-structured approach (core questions with room to probe) best fits your goal.

3.       Draft the questions:

·         Develop clear, open-ended questions aligned directly with your objectives.

·         Arrange them in a logical sequence, starting with easy “icebreakers” before moving to core technical or behavioral questions.

4.       Develop a scoring Rubrics: Create a standardized evaluation guide or rubric with clear criteria for scoring responses to minimize human bias and ensure fairness.

5.       Plan the Environment and Execution: Set a comfortable, distraction-free environment and determine the time limits for each section to ensure a smooth, professional process.

Uses of Interview

1.       Assessing Complex Skills: Evaluate higher-order abilities that written tests cannot easily measure, such as verbal communication, active listening, and interpersonal skills.

2.       Diagnosing Student understanding: Helps teachers uncover a learner’s thought process, identify specific misconceptions, and pinpoint learning gaps (formative assessment).

3.       Final Evaluation (Summative): Serves as a definitive grading tool or qualification check through formats like viva voces, oral exams and project defenses.

4.       Qualitative data collection: Acts as a vital tool in educational and action research to gather rich, deep insights into participant attitudes, experiences, and opinions.

5.       Selection and placement: Used by institutions and hiring panels to evaluate a candidate’s institutional fit, confidence and problem-solving abilities under pressure.

6.       Counselling and Guidance: Helps educators and counselors understand a student’s personal, emotional, or behavioral challenges in a safe, conversational space.



 Rubrics as an assessment tool

          A rubric is an assessment tool for communicating expectations of quality. It is designed to reflect the processes and outputs of learning, to support student self-reflection and self-assessment as well as communication between student, teacher and parents. It is usually in the form of a matrix with a list of indicators and a range of grading to rate the performance. A rubric provides a basis for self-evaluation, reflection, and pre-evaluation.

     Rubrics are generally thought to promote more consistent grading and to develop self-evaluation skills in students as they monitor their performance relative to the rubric.

Construction of Rubric tool

     The five essential rules for constructing a rubric:

1.       Limit to 3-5 Criteria: Focus only on the core learning objectives (e.g., Organization, Analysis) to keep the grading focused and manageable.

2.       Use a 4-level scale: An even number of performance levels (like 4-3-2-1) prevents “Middle-of-the-road” grading and forces a clear distinction between proficient and developing work.

3.       Write Bookends first: Draft the descriptors for the highest (“Excellent”) and lowest (“Unsatisfactory”) levels first, then build out the middle tiers.

4.       Use observable Actions: Avoid vague words like “good” or “poor”. Instead, describe concrete evidence (e.g., “Provides at least three supporting sources” vs. “Provides fewer than two sources”).

5.       Keep layout parallel: Ensure every performance tier evaluates the exact same elements in the same order so students can easily see how to improve.

Uses of Rubric as assessment tool

1.       Clarifies Expectations: It provides students with an explicit roadmap before they start, showing exactly what high-quality work looks like and how they will be evaluated.

2.       Speeds up grading: It streamlines the grading process for educators by allowing them to check off predetermined criteria rather than writing the same comments repeatedly.

3.       Ensures objectivity and consistency: It minimizes grader bias and grading drift, ensuring that the first paper and the last paper are evaluated against the exact same standard.

4.       Provides actionable feedback: It highlights a student’s specific strengths and weaknesses, showing them exactly where they succeeded and where hey need to improve.

5.       Facilitates Self-Assessment: It empowers students to review, evaluate, and edit their own work (or their peer’s work) against the criteria before final submission.

 

Comments