Treat the three studies as one evidence system
A baseline establishes the starting point, a midline checks progress and implementation assumptions, and an endline assesses results and lessons at or near completion. These studies are most useful when they are designed as connected stages rather than independent consulting assignments.
That connection requires consistent indicator definitions, comparable sampling logic, documented changes in tools and a clear record of contextual events that may affect interpretation. Without this continuity, apparent changes over time may reflect measurement differences rather than programme effects.
Start with the programme logic and decision needs
The study design should reflect the theory of change, results framework and the decisions managers or donors need to make. Not every indicator requires the same method, sample or frequency. Some outcomes are best measured through surveys, while others require administrative data, observation or qualitative inquiry.
A useful evaluation matrix can connect each question to indicators, data sources, methods and analytical approaches. This helps prevent questionnaires from becoming collections of items that are interesting but not tied to the assignment purpose.
Preserve comparability without becoming trapped by weak baseline design
Endline teams often inherit baseline tools with unclear questions, missing definitions or limited documentation. Repeating every weakness for the sake of comparability is not always the right solution. The evaluation should preserve comparable measures where they remain valid while documenting justified improvements.
Where direct comparison is impossible, analysts should explain the limitation transparently and use triangulation rather than forcing a misleading trend. Good evaluation protects credibility before it protects a preferred narrative.
Use qualitative evidence to explain change
Quantitative indicators can show whether an outcome changed, but they rarely explain why. KIIs, FGDs, IDIs, observation and Most Significant Change stories can identify mechanisms, implementation differences, unintended effects, barriers and contextual influences.
The strongest mixed-methods analysis does not present qualitative findings as a separate appendix to survey results. It integrates evidence around evaluation questions and uses each source to confirm, explain or challenge the others.
Plan for attribution realistically
Many development programmes cannot use experimental or quasi-experimental designs. In these settings, evaluators should avoid language that claims sole causality from simple before-and-after change. Contribution analysis, process evidence, triangulation, comparison with targets and careful examination of alternative explanations can support more defensible conclusions.
The appropriate strength of causal language should match the design. This is a sign of evaluation quality, not weakness.
Design recommendations for implementation
An endline is useful when its recommendations can be acted on. Recommendations should identify the problem, proposed action, responsible actor, priority and, where useful, timing. They should be traceable to evidence rather than introduced as generic good practice.
Management-response mechanisms can then track whether recommendations are accepted, adapted or rejected and why. This closes the loop between evaluation and organisational learning.
Context should be analysed, not treated as background
Programmes in Ethiopia may operate through drought, inflation, conflict, policy change, migration, market disruption or institutional turnover. These factors can influence both implementation and outcomes. Evaluation should therefore document major contextual changes and examine how they interact with the programme theory of change.
This does not mean attributing every weak result to context. It means distinguishing implementation performance from external conditions and explaining where the evidence supports each interpretation.
Data ownership and continuity should be planned
Baseline datasets, codebooks, tools, sampling documents and indicator calculations should be stored so that future evaluation teams can reuse them. Weak documentation creates avoidable problems at midline and endline, particularly when staff have changed or the original consultant is no longer available.
Clients can improve continuity by requiring clean datasets, syntax or calculation notes, final instruments and a concise methodological record as formal deliverables at each stage of the evaluation cycle.