Digital Forensics

Evidence review

What CFTT Testing Shows About Mobile Extraction Tools

Fourteen government test reports, read for what they measured — where extraction tools omitted data, where they reported it wrongly, and why the extraction method matters as much as the tool.

Published 6 September 2026  ·  Sources verified 12 September 2026  ·  14 min read

DEVI Digital Forensics  ·  ORCID iD 0009-0007-6471-1759

How to read this

Entry type
Evidence review
Sources verified
12 September 2026
Evidence classes
test, government

This entry reviews evidence published by other people. DEVI did not test any tool or examine any device for it. Blocks marked with an evidence class quote or summarize a source; sentences beginning “we” describe what DEVI did with those sources, which was to read them. Anything DEVI could not confirm is marked as such and left as an open question rather than written up as a finding. Sources change without notice, which is why the verification date is stated.

An extraction tool that reports nothing is easy to catch. An extraction tool that reports most of a phone, and quietly leaves out one application on one handset, is not — and that is the failure mode these test reports keep recording.

We read fourteen test reports published by the Department of Homeland Security's Science and Technology Directorate, covering eight vendors, and one test specification. This is what they say, what they do not say, and what an examiner can and cannot take from them.

We did not test any tool. Every measurement below belongs to the government laboratory that made it.

What CFTT is

The stated objective is "to provide measurable assurance to practitioners, researchers, and other applicable users that the tools used in computer forensics investigations provide accurate results," and the reports name the audience explicitly: developers improving tools, users making informed choices, and "the legal community and others to understand the tools' capabilities."

Tests run in the NIST CFTT laboratory. DHS S&T publishes the results. The governing document is the Mobile Device Forensic Tool Test Specification version 3.3, dated January 2025, and a self-service version of the same suite is available through Federated Testing.

We are not aware of another body of independent, published, measured evidence about these products. Everything else we found while looking is written by the vendors themselves.

How a tool gets tested

Each report opens with the same framing. The tool "was tested for its ability to acquire active data from the internal memory of supported mobile devices (i.e., Android, iOS)." Handsets are pre-populated with known data — contacts, messages, media, application data, network artifacts — and the tool's output is compared against what was planted. Anything missing, or present but wrong, is recorded as an anomaly.

Three things about that design decide how far a result travels, and the reports say all three out loud:

  • Every supported extraction technique is exercised. "Each supported data extraction technique was performed as well as each individual exploit across all devices."
  • Results are bounded by technique. "The data reported for the devices below varies based upon the data extraction technique supported. For instance, a physical and/or file system extraction will provide the user with more data than a logical extraction."
  • Anomalies are scoped, not general. "The anomalies listed below are specific to the phone type, OS and data extraction technique used."

That third sentence is the one to keep. It is the reason, stated by the laboratory itself, why no finding in this article can be read as a statement about a product in general.

What we read

Fourteen reports, listed here with the exact builds they cover, because the build is the claim:

All fourteen are indexed on the DHS S&T mobile device acquisition page, alongside reports we did not read. Separately, we also read Magnet AXIOM v1.2.1.6994, published 11 October 2018, for the connectivity-disruption claim discussed below. It is not part of the fourteen-report corpus counted elsewhere.

Social media is where the anomalies cluster

Every report we read that recorded anomalies at all recorded social media anomalies. The distinction the reports draw is between data "not reported" and data "partially reported," and partially reported almost always means account or profile information without message content.

Two things are worth separating there. Cellebrite's full file system run mostly degraded to partial reporting; Magnet's run returned nothing at all for three applications across eight iPhones. Those are different failures, and only the second one leaves an examiner with a report that looks complete because the application simply is not in it.

The extraction method is part of the result

This is the finding we would put first if an examiner read only one section.

Two reports cover the same Cellebrite acquisition build, UFED4PC 7.69.0.1397, paired with two different analysis products — Physical Analyzer 7.68.0.25 in one and Inseyets 10.2.101.352 in the other. Both ran iOS Advanced Logical on the Apple devices, and Advanced Logical, File System and File System APK Downgrade on the Android devices.

The two reports record identical anomaly lists. Social media data for Facebook, LinkedIn, Twitter/X, Instagram, SnapChat, WhatsApp and Pinterest "is not reported" on the OnePlus 10 Pro, Galaxy Note 10 and Galaxy Note 8. Contact photographs unreported on nine devices. MMS attachments unreported on four. Changing the analysis product changed nothing.

Set that against the Inseyets.UFED 10.3.0.260 report quoted above, which ran full file system extraction across largely the same handsets, and where the same applications mostly came back partial rather than absent.

We want to be careful about what that comparison licenses. The two sets of reports differ in both the acquisition build and the extraction method, so they do not isolate the method as the sole cause. What they do establish is narrower and still useful: swapping the analysis product moved nothing, and the report using a deeper acquisition method recorded far fewer outright omissions. A sentence of the form "Cellebrite reported X" is not a well-formed claim. The build and the extraction method have to be in it.

Signal and Telegram, and the method that produced that result

Two 2025–2026 reports are commonly cited for encrypted messengers, and both need their extraction method attached before they mean anything.

Now the part that changes the meaning. The Oxygen report's device table lists an extraction type per handset. Eight of its iPhones show "iTunes Backup, FFS." Seven show iTunes Backup alone — and the iPhone 16 Pro running iOS 18.1, the device the Telegram and Signal anomaly names, is one of those seven. No full file system extraction was performed on it. In the MD-LIVE report, every device in the test set was acquired logically; the Apple rows all read "Logical/iTunes Bkp."

So the honest statement is not that these tools cannot reach Signal and Telegram. It is that in the reports we read, the devices where Signal and Telegram went unreported were acquired by backup or logical methods, on the newest handsets in each set. That is consistent with the previous section rather than separate from it, and it is a much smaller claim than the one usually made from these reports.

When the connection drops

Both MSAB reports we read record the same behavior, in the same words, on both platforms.

The tail of that is important and is often dropped when this result is quoted. The examiner is not left with no signal at all — a terminal "Finished with errors" message appears. What does not appear is anything at the moment of the interruption, while the tool keeps running and keeps writing a partial extraction. The practical exposure is an examiner who reads the completion state rather than the error state, and an extraction that is short by an unknown amount.

Secondary sources have long said NIST recorded the same connectivity-disruption defect in Magnet Axiom in 2018. That report is public. We located it on the DHS publications path linked from the NIST CFTT mobile index and read it.

That is the same class of defect the 2025 MSAB reports record — no error at the moment of interruption — stated seven years earlier against a different product and build. It is not a finding about current Axiom; it is confirmation that the secondary claim about the 2018 report was accurate.

Reported, but wrong

Omissions are recoverable once you know about them. These are the anomalies where the output looks finished.

A value present in the wrong place, a name whose characters are reversed, a PDF that is listed but will not open — none of these produce an empty cell that prompts a second look.

The clean reports, and how far they reach

Two Graykey reports recorded no anomalies at all.

Scope belongs next to that. The 2025 Graykey and Axiom report covers twelve devices: eleven Apple handsets running iOS 15.1, 15.6 or 17.4.1, and one Samsung Galaxy S22. The 2023 GrayKey Android report covers four devices — Galaxy S22, Galaxy Z Fold3 5G, Galaxy Tab S8 and Pixel 4 — and it did record anomalies, including "Notes were not reported. (Device: Google Pixel 4)" and three social media anomalies.

These are the smallest test sets among the reports we read. A clean report on twelve devices is a real result on twelve devices. It is not evidence about a thirteenth.

Two products we did not find

Cellebrite Premium and Magnet Verakey are marketed on device access and consent-based extraction respectively. We looked for test reports covering them.

We searched the DHS S&T mobile device acquisition index and the NIST CFTT mobile index on 6 September 2026. Neither page's text contains the word "Premium" or the word "Verakey." Neither product appears in any of the fourteen reports we read; Cellebrite is covered through UFED, UFED4PC, Inseyets and Physical Analyzer, and Grayshift through the two Graykey reports.

We are stating an absence from two index pages and fourteen reports on one date. We did not open every report linked from either index, and we make no claim about validation that may exist elsewhere, unpublished, or under another name.

What the dates say, and what they do not

The Cellebrite Inseyets report published on 19 May 2025 carries "September 2024" on its cover. That is roughly an eight-month gap between testing and publication, and it is visible in the hardware: the Apple devices in these test sets run iOS 15.1, 15.6, 16.7, 16.7.3 and 17.4.1, with iOS 18.1 appearing on a single handset in one report.

Counting test-result documents linked from the DHS S&T index on 6 September 2026, by the year of publication:

  • 2022 — three
  • 2023 — seventeen
  • 2024 — two
  • 2025 — fourteen
  • 2026, through August — one

We are reporting counts, not a trend. CFTT publishes in irregular batches — four reports share a single date in March 2025 — the 2024 figure was followed by fourteen in 2025, and the test specification was updated to version 3.3 in January 2025. No agency we read has stated that testing has slowed or stopped, and we did not find evidence that it has. The counts are here so a reader can see what the published record looks like, and to make one narrower point: independent results describe builds that are, routinely, a year or more behind what is currently sold.

What a CFTT report establishes

Taken together, and stated as carefully as we can:

  • A result is version-specific. It describes the build named on the cover, not the product line.
  • A result is device-specific. It describes the handsets in the table, at the operating system versions in the table.
  • A result is method-specific. The reports say so themselves.
  • A result is bounded by the test corpus. An application that was not populated on a device produces "NA," not a finding.
  • A clean report does not establish that other versions, devices or methods are clean. An anomalous report does not establish that a product is unreliable in general.

What we could not verify

Not independently verified

Whether any independent validation of Cellebrite Premium or Magnet Verakey exists outside the two index pages and fourteen reports we searched. The 2018 Magnet AXIOM report cited above is a separate historical CFTT result; it is not a Premium or Verakey validation.

Not independently verified

We did not reconcile the report counts published on the NIST CFTT index against those on the DHS S&T index; the two pages link different numbers of documents. Our counts above are from the DHS index only.

Sources and evidence

Sources verified on 12 September 2026

Vendor
CellebriteMagnet ForensicsGrayshiftMSABOxygen ForensicsBelkasoftElcomSoftGMDSOFT
Product and version
Cellebrite Inseyets.UFED 10.3.0.260Cellebrite Inseyets.PA 10.3.0.3169Cellebrite UFED4PC 7.69.0.1397Cellebrite Physical Analyzer 7.68.0.25Cellebrite Inseyets 10.2.101.352Magnet Graykey AppLogic 4.1.0.27711255Magnet Axiom Examine 8.0.0.39753Magnet Axiom 8.1.0.42087Magnet AXIOM v1.2.1.6994Graykey OS 1.7.3.19461530GrayKey Advanced Rev D OS 1.10.1.21812323MSAB XRY Kiosk 10.9.0MSAB XRY Office 10.9.0MSAB XAMN 7.9.0Oxygen Forensic Detective 17.1.0.131Belkasoft Evidence Center X 2.7.19314Belkasoft Evidence Center X 2.5.16519Elcomsoft iOS Forensic Toolkit 8.30GMDSOFT MD-LIVE 3.6.14.1440
Platform
iOSAndroid
Operating system
iOS 15.1iOS 15.6iOS 16.7iOS 16.7.3iOS 17.4.1iOS 18.1Android 8.1.0Android 9Android 10Android 11Android 12Android 13Android 14
Application
SignalTelegramWhatsAppFacebookInstagramTwitterSnapchatLinkedInPinterestTikTokRedditDiscord
Source
NIST Computer Forensics Tool Testing Program (CFTT)DHS Science and Technology DirectorateNational Institute of Justice
Report reviewed
25_0519 Cellebrite Inseyets.UFED 10.3.0.260 / Inseyets.PA 10.3.0.316925_0519 Cellebrite UFED4PC 7.69.0.1397 / Inseyets 10.2.101.35225_0424 Cellebrite UFED4PC 7.69.0.1397 / PA 7.68.0.2525_0324 Magnet Graykey AppLogic 4.1.0 / Axiom Examine 8.0.025_0324 Magnet Axiom 8.1.0.4208718_1011 Magnet AXIOM v1.2.1.699425_0424 MSAB XRY Kiosk 10.9.025_0324 MSAB XRY Office 10.9.0 / XAMN 7.9.025_0501 Oxygen Forensic Detective 17.1.0.13125_0805 Belkasoft Evidence Center X 2.7.1931425_0527 Belkasoft Evidence Center X 2.5.1651924_0513 Elcomsoft iOS Forensic Toolkit 8.3023_0411 GrayKey Advanced Rev D OS 1.10.122_0615 Graykey OS 1.7.3.1946153026_0828 GMDSOFT MD-LIVE 3.6.14.1440Mobile Device Forensic Tool Test Specification v3.3

Indexed under

  • Android
  • iOS
  • Tool Validation
  • Government Testing
  • Mobile Extraction
  • Cellebrite
  • Examiner Workflow

Independent and educational. DEVI Digital Forensics is an independent educational project created by digital forensic practitioners outside of their official employment. It is not sponsored, reviewed, approved, or endorsed by any contributor's employing agency.

Forensic behavior changes between operating system versions, application versions, extraction methods, and tool versions. Validate every finding against your own data, and do not interpret an artifact in isolation. Read the full statement and methodology.

← All findings