Insect bioassays

Support guide for design, analysis and the Petri-dish practical

Author

Dr Joe Roberts

Published

July 2026

Use the four decisions

Decision Question to answer before data collection
Define What contrast, behaviour, population and conditions does the hypothesis specify?
Match Does the arena allow that behaviour, and does the endpoint measure it?
Protect What is independently treated, and what other variation must be controlled or blocked?
Claim Which analysis estimates the planned contrast, and where must the conclusion stop?

The live presentation uses one evidence ladder:

  1. Stimulus received: what reaches the insect?
  2. Detection: can the sensory system register it?
  3. Behaviour: what does the insect do?
  4. Consequence: what changes afterwards?

A result supports the rung that was measured. Later rungs need separate evidence (Roberts et al., 2023).

Match the method to the claim

Claim Suitable starting method Example endpoint Main limitation
Detection Electroantennography (EAG), gas chromatography–electroantennographic detection (GC–EAD) or single-sensillum recording Electrical response Detection is not preference. A negative whole-antenna response does not prove that no receptor detects the cue (Jacob, 2018).
Relative orientation Y-tube or four-arm olfactometer First arm entered or time in odour fields The choice is relative to the alternatives and depends on the delivered field.
Directed movement Wind tunnel or tracked gradient arena Upwind displacement, heading or plume encounters Throughput is lower and plume structure must be measured.
Acceptance Leaf-disc, Petri-dish or intact-plant assay Settling, contact, departure or feeding initiation Confinement and detached tissue can change behaviour.
Feeding process Electrical penetration graph Probing and feeding waveform durations Handling, wiring and tethering can alter normal movement and feeding (Walker et al., 2024).
Biological consequence Cage, whole-plant or semi-field assay Survival, fecundity, colonisation or crop outcome Environmental control falls as realism increases.

Host use is not one outcome

Host use can contain at least four linked stages:

  1. habitat location;
  2. host location;
  3. host acceptance;
  4. host suitability.

One assay rarely establishes the full sequence. For example, Y-tube orientation does not establish feeding success or population growth (Powell et al., 2006).

Protect the comparison

Controls and validation

For every control, state the alternative explanation it addresses.

Check Question
Matched comparison Are solvent, volume, evaporation, handling and presentation identical apart from the intended treatment?
Identical blank Does the arena itself create a position bias?
Stimulus field Is airflow, light, humidity, temperature or surface loading within its acceptance range?
Carry-over Does the cleaning interval prevent residues and order effects?
Measurement Are cameras, zones, clocks and scoring rules validated?
Positive control Can the assay detect a known response today, where a reliable positive control exists?

A positive control distinguishes no treatment effect from no functioning assay. It is not available for every species or endpoint (Roberts et al., 2023).

Handling, recovery and acclimation

  • Use one defined transfer method.
  • Avoid CO₂ or chilling unless their effect on the endpoint has been tested.
  • Pre-specify recovery and arena acclimation.
  • Pilot the interval using blank trials until baseline behaviour is stable.
  • Record the tool, exposure duration, time off host, recovery, temperature, light and operator.

There is no universal acclimation period. Transfer and dislodgement can themselves trigger immobility or escape-like movement (Johnson, 1958; Phelan et al., 1976; Powell and Bale, 2006). CO₂ anaesthesia can also impair later motor behaviour, so recovery must be demonstrated for the species and endpoint rather than assumed (Bartholomew et al., 2015).

Experimental unit and hierarchy

The experimental unit is the smallest unit independently assigned to a treatment. The observational unit is where a response is recorded.

Twenty aphids placed together in one treated dish provide twenty observations but only one independent treatment assignment. A mixed model can represent real grouping; it cannot manufacture replication that the experiment never created (Percie du Sert et al., 2020).

In the presentation’s illustrative example, the treated dish counts were 19, 20, 14, 20 and 8 responses out of 20; the control dish counts were 17, 13, 8, 3 and 18 out of 20. Pooling gives 81/100 versus 59/100 and a two-sided pooled z-test of proportions gives (p=0.000687), but that test incorrectly treats 200 aphids as independent. The valid demonstration first calculates the ten dish proportions. The observed absolute difference between the two group means is 0.22. Of all (=252) allocations of those ten dish values, 66 are at least as extreme, giving an exact two-sided permutation (p=66/252=0.262).

Plan sample size and non-response

Pilot first

Use the pilot to estimate:

  • baseline response and non-response rates;
  • variance among independently treated units;
  • technical-failure rate;
  • trials per hour and realistic block size;
  • whether the positive control and blank meet pre-specified criteria.

Choose a smallest effect that would change the biological conclusion. Base the sample-size justification on that effect, the experimental unit and the planned analysis. For hierarchical or time-to-event designs, simulation is often more transparent than a closed-form calculation (Lakens, 2022).

The dish-level example from the presentation

For an unblocked 5-versus-5 exact randomisation test using the absolute difference in group means, there are only (=252) allocations. The minimum attainable two-sided p-value is therefore (2/252=0.0079). That is a design constraint, not a defect in the analysis.

The pooled within-treatment standard deviation of the ten illustrative dish proportions is 0.289. Under a simple two-group normal approximation with a two-sided 5% significance level, 80% power and a 0.25 absolute difference, (2(z_{0.975}+z_{0.80})2(0.289)2/(0.25)^2=21.0), which rounds up to 22 independent dishes per treatment. This calculation is a bridge to sample size planning, not a reusable magic number: justify the smallest worthwhile effect, use pilot variation at the experimental-unit level, allow for technical losses and simulate the model that will actually be fitted.

Keep non-response in the question

Before data collection define:

  • what counts as a response;
  • how an insect reaching the time limit is represented;
  • which events are biological outcomes;
  • which equipment or handling failures justify exclusion.

Report released, responding and excluded numbers with reasons by treatment. For latency, a valid non-responder at the time limit is normally right-censored. Removing non-responders can change the question being estimated.

Diagnostic sequence for high non-response

  1. Delivery: did the cue reach the insect?
  2. State: was the insect eligible, recovered and acclimated?
  3. Arena: did geometry or confinement suppress the behaviour?
  4. Endpoint: did the scoring rule or observation window miss the response?
  5. Assay function: do blanks and a validated positive control behave as expected?

Increase replication only after deciding whether the problem is technical or biological.

Choose an analysis family

This table is a starting map, not a substitute for checking assumptions.

Endpoint Starting analysis Features that must be represented
Yes/no response Binomial logistic regression Blocks, repeated units, shared arenas and planned interactions
A/B/non-response or drop/walk/remain Multinomial regression, or a pre-specified two-part analysis Sparse categories, clustering and a clear reference category
Time until response Survival model Right-censoring at the time limit, blocks and shared trials
Counts or rates Poisson or negative-binomial regression Observation-time offset, extra variation and grouping
Continuous summary Linear model or mixed model Repeated observations, blocks and distributional fit
Repeated movement trajectory Mixed model or generalised additive mixed model Autocorrelation and repeated records from each insect or arena
Behavioural state sequence Hidden Markov model State definition, transition probabilities, temporal dependence and validation (Langrock et al., 2012)
Time budget Compositional analysis Components sum to a fixed total and are not independent

Terms in plain language

  • GLM: a regression model for outcomes such as yes/no or counts.
  • GLMM: a GLM with supported random effects for real grouping or repeated units.
  • Offset: represents unequal exposure or observation time in a rate model.
  • Right-censoring: the event was not observed before the study’s time limit, but the available time information is retained.
  • Hidden Markov model: estimates unobserved behavioural states and the probability of moving between them from a time sequence.

Report the experimental-unit sample size, planned contrast, effect size with a 95% confidence interval and relevant diagnostic checks (Bolker et al., 2009).

Preregister and check permissions

Preregister the decisions that can drift

Before data collection, time-stamp:

  • the primary endpoint and claim;
  • the experimental unit and allocation;
  • sample-size and stopping rules;
  • non-response and exclusion rules;
  • planned model, contrasts and secondary analyses.

OSF Registrations can be public or embargoed. Record and explain deviations. The ARRIVE Study Plan is a useful planning checklist for living invertebrate work (Nosek et al., 2018; Percie du Sert et al., 2020).

Petri-dish departure practical

Fixed hypothesis

A standardised brush touch increases the probability that a settled aphid leaves a field-bean leaf disc within 60 seconds, compared with an approach-only control, when the primary touch effect is averaged equally across the two supplied clones.

Worked protocol after the class agrees its design

  1. Use one named clone of pea aphid and one named clone of black bean aphid. Select wingless adults within a defined age window. A clone is a genetically matched lineage; one clone per species does not support a general species comparison.
  2. Place a field-bean leaf disc flat on damp filter paper in the base of a 90-mm Petri dish. Standardise leaf age, position, disc surface and preparation time, or include their source in the blocking scheme.
  3. Place one aphid at the disc centre. Apply a pre-specified settling criterion and maximum settling period. Record non-settlers.
  4. For contact, use one light one-second touch to a defined body region with one fine-brush bristle. Withdraw when the bristle just flexes.
  5. For the no-contact control, reproduce the approach and one-second pause but stop 2 mm above the aphid. Do not repeat either stimulus.
  6. Film for 60 seconds. Record temperature, light, time and operator.
  7. Use every aphid, leaf disc and dish once. Randomise the four clone-by-treatment combinations within a balanced complete block.

Each complete block contains four dishes, one for every clone-by-treatment combination. With (B) complete blocks, the pooled class contains (4B) dishes in total, (2B) dishes per treatment and (B) dishes per clone-by-treatment cell. A block is one short run made under common conditions, normally by one group or operator. Treat the exercise as a pilot for estimating responses and validating the workflow, not automatically as a definitive treatment-by-clone test.

Outcomes

  • Primary: left the disc within 60 seconds, yes or no.
  • Secondary: departure latency and any pre-specified movement descriptor.
  • Technical failures: accidental displacement, injury or lost recording, reported separately.

A flat-disc arena measures departure from the disc; it cannot reproduce or identify the natural dropping distance seen on intact plants. Do not claim a dropping response from this design. If dropping itself is the endpoint, use an elevated intact-leaf or whole-plant arena designed to capture it (Braendle and Weisser, 2001).

Analysis and claim

For the pooled class exercise, start with a binomial generalised linear model for leaves/remains, with treatment, clone and complete block fitted as fixed effects. Report model-predicted departure probabilities and their treatment contrast—preferably an absolute risk difference—with a 95% confidence interval. Treat the treatment-by-clone interaction as exploratory unless the number of complete blocks supports it. In a larger study with many sampled blocks, a pre-specified random-intercept model may be more appropriate.

The defensible conclusion concerns the two tested clones, this standardised tactile cue and the recorded flat-disc conditions. It does not establish a general species difference, a natural dropping response or a response to a live predator (Braendle and Weisser, 2001; Gish, 2021; Rasekh et al., 2010).

Debrief for the flawed-protocol activity

The scenario contains more than six defensible faults:

  1. mixed age and state;
  2. direct transfer without recovery or acclimation;
  3. treatment solvent not matched on the control side;
  4. no defined evaporation interval;
  5. treatment always presented on the left;
  6. twenty aphids share a treatment assignment;
  7. dish-level replication is only five;
  8. centre aphids are removed informatively;
  9. the one-hour endpoint does not establish repellency;
  10. one cohort, day and extract batch limit generality;
  11. no validated known-response control despite one being available;
  12. analysis treats subsamples as independent;
  13. the claim exceeds the measured endpoint.

The repair is not one statistical model. Align the comparison, handling, experimental unit, endpoint, validation and claim before choosing the model.

References

Bartholomew, Burdett, VandenBrooks, Quinlan, and Call (2015). Impaired climbing and flight behaviour in drosophila melanogaster following carbon dioxide anaesthesia. Scientific Reports. https://doi.org/10.1038/srep15298.
Bolker et al. (2009). Generalized linear mixed models: A practical guide for ecology and evolution. Trends in Ecology & Evolution. https://doi.org/10.1016/j.tree.2008.10.008.
Braendle and Weisser (2001). Variation in escape behavior of red and green clones of the pea aphid. Journal of Insect Behavior. https://doi.org/10.1023/A:1011124122873.
Defra (2024). Technical assessment of scientific authorisation applications. https://planthealthportal.defra.gov.uk/scientific-authorisations/scientific-authorisation-guidance-for-authorisation-holders-and-applicants/the-scientific-authorisation-process-in-england-and-wales/technical-assessment-of-scientific-authorisation-applications/.
Gish (2021). Aphids detect approaching predators using plant-borne vibrations and visual cues. Journal of Pest Science. https://doi.org/10.1007/s10340-020-01323-6.
Home Office (2024). The operation of the animals (scientific procedures) act 1986. https://www.gov.uk/government/publications/the-operation-of-the-animals-scientific-procedures-act-1986/the-operation-of-the-animals-scientific-procedures-act-1986-aspa-accessible.
Jacob (2018). Current source density analysis of electroantennogram recordings: A tool for mapping the olfactory response in an insect antenna. Frontiers in Cellular Neuroscience. https://doi.org/10.3389/fncel.2018.00287.
Johnson (1958). Factors affecting the locomotor and settling responses of alate aphids. Animal Behaviour. https://doi.org/10.1016/0003-3472(58)90004-6.
Lakens (2022). Sample size justification. Collabra: Psychology. https://doi.org/10.1525/collabra.33267.
Langrock, King, Matthiopoulos, Thomas, Fortin, and Morales (2012). Flexible and practical modeling of animal telemetry data: Hidden markov models and extensions. Ecology. https://doi.org/10.1890/11-2241.1.
Nosek, Ebersole, DeHaven, and Mellor (2018). The preregistration revolution. Proceedings of the National Academy of Sciences. https://doi.org/10.1073/pnas.1708274114.
Percie du Sert et al. (2020). Reporting animal research: Explanation and elaboration for the ARRIVE guidelines 2.0. PLOS Biology. https://doi.org/10.1371/journal.pbio.3000411.
Phelan et al. (1976). Orientation and locomotion of apterous aphids dislodged from their hosts by alarm pheromone. Annals of the Entomological Society of America. https://doi.org/10.1093/aesa/69.6.1153.
Powell and Bale (2006). Effect of long-term and rapid cold hardening on the cold torpor temperature of an aphid. Physiological Entomology. https://doi.org/10.1111/j.1365-3032.2006.00527.x.
Powell, Tosh, and Hardie (2006). Host plant selection by aphids: Behavioural, evolutionary, and applied perspectives. Annual Review of Entomology. https://doi.org/10.1146/annurev.ento.51.110104.151107.
Rasekh et al. (2010). Ant mimicry by an aphid parasitoid, lysiphlebus fabarum. Journal of Insect Science. https://doi.org/10.1673/031.010.12601.
Roberts, Clunie, Leather, Harris, and Pope (2023). Scents and sensibility: Best practice in insect olfactometer bioassays. Entomologia Experimentalis et Applicata. https://doi.org/10.1111/eea.13351.
UK Parliament (2022). Animal welfare (sentience) act 2022. https://www.legislation.gov.uk/ukpga/2022/22/section/5.
Walker, Fereres, and Tjallingii (2024). Guidelines for conducting, analyzing, and interpreting electrical penetration graph (EPG) experiments on herbivorous piercing-sucking insects. Entomologia Experimentalis et Applicata. https://doi.org/10.1111/eea.13434.