Look closely enough and there is an entire world moving beneath our feet, over our crops, and through our forests, much of it smaller than a fingernail, and much of it still unknown. Pakistan’s insects occupy almost every ecological niche imaginable, whether it’s the grasshoppers of the country’s arid plains or the butterflies and moths of its forests, the aphids feeding on its crops, the beetles breaking down dead matter or the mosquitoes that can carry disease. Traditionally, studying an insect meant placing it under a microscope, examining its physical characteristics and comparing it with existing descriptions. Today, that process is being augmented by DNA barcodes, genomic databases, machine learning, and artificial intelligence. The question has evolved from What insect is this? to What can its genes, images and data tell us that our eyes cannot? Two species can look alike or a damaged or immature specimen may not have the necessary characteristics to help identify it. When physical traits reach their diagnostic limit, bioinformatics steps in to read the unalterable ledger coded inside the cells—the insect’s DNA. DNA barcoding uses a standard DNA sequence to identify species by comparing with reference records. Illustration by LarissaFruehe via Wikimedia Commons, licensed under CC BY-SA 4. 0. Pakistan has already demonstrated what can be achieved at a national level through this method. A DNA-barcode survey of insect biodiversity examined material collected from 1, 858 sites between 2010 and 2019. It eventually generated 50, 592 barcode records, 49, 363 of which were assigned to 6, 590 Barcode Index Numbers (BINs). Simply put, a BIN is a reference code used in the global Barcode of Life Data System (BOLD) database. When researchers analyse a short piece of an insect’s DNA, BOLD compares it with existing records and clusters similar specimens into the same group under a specific BIN code. This makes it possible to group insects that likely belong to the same species, even when their formal scientific name is unknown. This is how the BIN serves as a practical working proxy name for a species. The study produced Pakistan’s first comprehensive DNA barcode profile of its insects, and the first survey of its kind in South Asia. But perhaps the most interesting finding was not what the researchers identified, but how much they couldn’t. Only 21 per cent of the BINs could be matched to species that already had a name in global databases. At the time of the analysis, 59pc of the BINs had been recorded in BOLD, only from Pakistan. That does not necessarily mean these insects exist nowhere else. Much of Pakistan, as well as neighbouring countries, has simply not been sampled evenly. In other words, we are still missing large pieces of the insect puzzle. A barcode library can help make those gaps visible. It can show researchers where they need to look next, collect more specimens and, crucially, carry out the painstaking work of taxonomy—identifying, classifying and naming the species. And with researchers sharing data across borders, some of these unknowns may finally begin to get names. The Pakistan insect DNA-barcode survey lists material collected from 1, 858 sites between 2010 and 2019. The final dataset has 50, 592 barcode records, with 49, 363 specimens assigned to 6, 590 BINs. Source: “A DNA barcode survey of insect biodiversity in Pakistan” — Muhammad Ashfaq et al. (2022). Looking alike, behaving differently Bioinformatics, which is essentially the use of science and math to store, analyse, and understand large sets of biological data, is crucial to deploy if the insect is an agricultural pest. So, for example, the whitefly Bemisia tabaci looks like one little white fly but scientists know that there is a complex of closely related species that are visually similar. These groups differ in how they spread plant viruses and respond to insecticides that seek to kill the bugs. In one Pakistani whitefly study, specimens were collected from 255 locations in Punjab and Sindh and 173 records were added. The analysis resulted in the division of the material into 15 BINs, the identification of several putative species, and the identification of a previously unknown genetic group labelled “Pakistan”. Even though they may look alike, significant differences may be overlooked without DNA analysis. A later study went beyond a short barcode and compared the whole genome of the Asia II 1 whitefly with a MEAM1 (Middle East-Asia Minor 1) reference genome. MEAM1 is a highly invasive biotype of the sweetpotato whitefly ( Bemisia tabaci ). It is globally recognised as one of the most destructive agricultural pests due to its ability to transmit crop viruses and rapidly evolve resistance to chemicals. Whitefly, Bemisia tabaci adults on a watermelon leaf. Credit: CSIRO via Wikimedia Commons, licensed under CC BY 3. 0. Unmodified. Researchers compared the genome of the Asia II 1 whitefly with the MEAM1 whitefly, using the latter as a genetic baseline. The comparison revealed about 2. 33 million single-letter differences in the DNA, known as SNPs, along with another 202, 479 insertions and deletions. They identified variants in 14 genes previously associated with potential insecticide resistance. This is where bioinformatics comes in handy. Instead of trying to make sense of millions of genetic differences individually, computers can help researchers narrow the list down to a much smaller group of candidates worth looking into. But a machine can only take the investigation so far. A genetic difference may look important on a screen without actually making a whitefly better at surviving an insecticide. That has to be tested in the laboratory and, ultimately, in the field. A separate study tested whiteflies collected from five districts of Punjab between 2017 and 2019. Their response differed according to the location and insecticide tested. The lesson is simple. The computer can point researchers towards the suspects, but experiments establish whether they are actually guilty. Understanding insecticide resistance therefore requires all three pieces: bioinformatics to find the clues, laboratory experiments to test them, and regular field monitoring to see what is happening in real populations. Adult pink bollworm moth, Pectinophora gossypiella. Credit: Mississippi State University Archive, Mississippi State University, Bugwood. org, via Wikimedia Commons; licensed under CC BY 3. 0 US. Image No. 1265079. Unmodified. Bioinformatics is also being used to research the pink bollworm, another serious cotton pest. Scientists have used thousands of genetic markers to compare populations, including specimens from Punjab and Sindh. Another study used records from 17 districts of Punjab, laboratory observations and temperature-based models to estimate how infestations might change in the future. These predictions may guide researchers on where to concentrate their efforts, but they are not definite predictions. Field conditions such as weather, variety, planting time, and pest management can all have an impact. What Pakistan can learn from global projects The Anopheles gambiae 1000 Genomes Project (Ag1000G) is a collection of mosquito genome data in several countries across sub-Saharan Africa. This common information is used by researchers to monitor genetic variation, mosquito population movement and the spread of insecticide resistance. A similar network of universities, research institutes and provincial field teams is needed in Pakistan. This would enable the researchers to compare an insect collected in Faisalabad with the insects of Sindh, Khyber Pakhtunkhwa or Balochistan. Blood-feeding Anopheles gambiae mosquitoes. Credit: Johns Hopkins Malaria Research Institute via PLOS Biology and Wikimedia Commons; licensed under CC BY 2. 5. Unmodified One example is the i5K initiative, which seeks to provide detailed genetic information for thousands of insects and related species. Pakistan can start with the key crop pests, disease vectors, pollinators and other beneficial insects. Every DNA record should be associated with a well-identified and curated specimen, and with basic information on its location and time of collection, and the host plant or animal species from which it was obtained. This would make it easier to review, compare and reuse information by other researchers. Insects can, of course, also be misidentified, leading to the subsequent misidentification of their DNA record, which is why scientists don’t solely rely on bioinformatics. Contamination, incomplete databases, small samples and missing collection details can lead even the best astray. This is when entomologists exit the lab and descend into the field. Bioinformatics cannot be a substitute for fieldwork, taxonomy, or lab experiments. Rather, it can unite them. It can reveal genetic differences between insects that look alike, highlight genes worth studying, and show how insect populations differ. Pakistan has already taken a good beginning. The next step is to enhance reference collections, exchange information, and recruit researchers in both entomology and computer-based analysis, and continue field sampling. To look at an insect today is to see a dual identity: a living organism birthed by local crops and changing climates, and a precise record written in DNA. Pakistan’s scientific future depends on mastering both, learning to read the digital blueprint inside the cell without ever losing touch with the insect in the field. Header image by Mohsin Alam
The DNA detectives are changing the game for insect research in Pakistan
RELATED ARTICLES



