How to read this
- Entry type
- Evidence review
- Sources verified
- 12 September 2026
- Evidence classes
- test, government
This entry reviews evidence published by other people. DEVI did not test any tool or examine any device for it. Blocks marked with an evidence class quote or summarize a source; sentences beginning “we” describe what DEVI did with those sources, which was to read them. Anything DEVI could not confirm is marked as such and left as an open question rather than written up as a finding. Sources change without notice, which is why the verification date is stated.
An extraction tool that reports nothing is easy to catch. An extraction tool that reports most of a phone, and quietly leaves out one application on one handset, is not — and that is the failure mode these test reports keep recording.
We read fourteen test reports published by the Department of Homeland Security's Science and Technology Directorate, covering eight vendors, and one test specification. This is what they say, what they do not say, and what an examiner can and cannot take from them.
We did not test any tool. Every measurement below belongs to the government laboratory that made it.
What CFTT is
The stated objective is "to provide measurable assurance to practitioners, researchers, and other applicable users that the tools used in computer forensics investigations provide accurate results," and the reports name the audience explicitly: developers improving tools, users making informed choices, and "the legal community and others to understand the tools' capabilities."
Tests run in the NIST CFTT laboratory. DHS S&T publishes the results. The governing document is the Mobile Device Forensic Tool Test Specification version 3.3, dated January 2025, and a self-service version of the same suite is available through Federated Testing.
We are not aware of another body of independent, published, measured evidence about these products. Everything else we found while looking is written by the vendors themselves.
How a tool gets tested
Each report opens with the same framing. The tool "was tested for its ability to acquire active data from the internal memory of supported mobile devices (i.e., Android, iOS)." Handsets are pre-populated with known data — contacts, messages, media, application data, network artifacts — and the tool's output is compared against what was planted. Anything missing, or present but wrong, is recorded as an anomaly.
Three things about that design decide how far a result travels, and the reports say all three out loud:
- Every supported extraction technique is exercised. "Each supported data extraction technique was performed as well as each individual exploit across all devices."
- Results are bounded by technique. "The data reported for the devices below varies based upon the data extraction technique supported. For instance, a physical and/or file system extraction will provide the user with more data than a logical extraction."
- Anomalies are scoped, not general. "The anomalies listed below are specific to the phone type, OS and data extraction technique used."
That third sentence is the one to keep. It is the reason, stated by the laboratory itself, why no finding in this article can be read as a statement about a product in general.
What we read
Fourteen reports, listed here with the exact builds they cover, because the build is the claim:
- Cellebrite Inseyets.UFED 10.3.0.260 / Inseyets.PA 10.3.0.3169, published 19 May 2025
- Cellebrite UFED4PC 7.69.0.1397 / Inseyets 10.2.101.352, published 19 May 2025
- Cellebrite UFED4PC 7.69.0.1397 / Physical Analyzer 7.68.0.25, published 24 April 2025
- Magnet Graykey AppLogic 4.1.0.27711255 / Axiom Examine 8.0.0.39753, published 24 March 2025
- Magnet Axiom 8.1.0.42087, published 24 March 2025
- MSAB XRY Kiosk 10.9.0, published 24 April 2025
- MSAB XRY Office 10.9.0 / XAMN 7.9.0, published 24 March 2025
- Oxygen Forensic Detective 17.1.0.131, published 1 May 2025
- Belkasoft Evidence Center X 2.7.19314, published 5 August 2025
- Belkasoft Evidence Center X 2.5.16519, published 27 May 2025
- Elcomsoft iOS Forensic Toolkit 8.30, published 13 May 2024
- GrayKey Advanced Rev D, OS 1.10.1.21812323, AppLogic 3.6.0, published 11 April 2023
- Graykey OS 1.7.3.19461530, published 15 June 2022
- GMDSOFT MD-LIVE 3.6.14.1440, published 28 August 2026
All fourteen are indexed on the DHS S&T mobile device acquisition page, alongside reports we did not read. Separately, we also read Magnet AXIOM v1.2.1.6994, published 11 October 2018, for the connectivity-disruption claim discussed below. It is not part of the fourteen-report corpus counted elsewhere.
Social media is where the anomalies cluster
Every report we read that recorded anomalies at all recorded social media anomalies. The distinction the reports draw is between data "not reported" and data "partially reported," and partially reported almost always means account or profile information without message content.
Two things are worth separating there. Cellebrite's full file system run mostly degraded to partial reporting; Magnet's run returned nothing at all for three applications across eight iPhones. Those are different failures, and only the second one leaves an examiner with a report that looks complete because the application simply is not in it.
The extraction method is part of the result
This is the finding we would put first if an examiner read only one section.
Two reports cover the same Cellebrite acquisition build, UFED4PC 7.69.0.1397, paired with two different analysis products — Physical Analyzer 7.68.0.25 in one and Inseyets 10.2.101.352 in the other. Both ran iOS Advanced Logical on the Apple devices, and Advanced Logical, File System and File System APK Downgrade on the Android devices.
The two reports record identical anomaly lists. Social media data for Facebook, LinkedIn, Twitter/X, Instagram, SnapChat, WhatsApp and Pinterest "is not reported" on the OnePlus 10 Pro, Galaxy Note 10 and Galaxy Note 8. Contact photographs unreported on nine devices. MMS attachments unreported on four. Changing the analysis product changed nothing.
Set that against the Inseyets.UFED 10.3.0.260 report quoted above, which ran full file system extraction across largely the same handsets, and where the same applications mostly came back partial rather than absent.
We want to be careful about what that comparison licenses. The two sets of reports differ in both the acquisition build and the extraction method, so they do not isolate the method as the sole cause. What they do establish is narrower and still useful: swapping the analysis product moved nothing, and the report using a deeper acquisition method recorded far fewer outright omissions. A sentence of the form "Cellebrite reported X" is not a well-formed claim. The build and the extraction method have to be in it.
Signal and Telegram, and the method that produced that result
Two 2025–2026 reports are commonly cited for encrypted messengers, and both need their extraction method attached before they mean anything.
Now the part that changes the meaning. The Oxygen report's device table lists an extraction type per handset. Eight of its iPhones show "iTunes Backup, FFS." Seven show iTunes Backup alone — and the iPhone 16 Pro running iOS 18.1, the device the Telegram and Signal anomaly names, is one of those seven. No full file system extraction was performed on it. In the MD-LIVE report, every device in the test set was acquired logically; the Apple rows all read "Logical/iTunes Bkp."
So the honest statement is not that these tools cannot reach Signal and Telegram. It is that in the reports we read, the devices where Signal and Telegram went unreported were acquired by backup or logical methods, on the newest handsets in each set. That is consistent with the previous section rather than separate from it, and it is a much smaller claim than the one usually made from these reports.
When the connection drops
Both MSAB reports we read record the same behavior, in the same words, on both platforms.
The tail of that is important and is often dropped when this result is quoted. The examiner is not left with no signal at all — a terminal "Finished with errors" message appears. What does not appear is anything at the moment of the interruption, while the tool keeps running and keeps writing a partial extraction. The practical exposure is an examiner who reads the completion state rather than the error state, and an extraction that is short by an unknown amount.
Secondary sources have long said NIST recorded the same connectivity-disruption defect in Magnet Axiom in 2018. That report is public. We located it on the DHS publications path linked from the NIST CFTT mobile index and read it.
That is the same class of defect the 2025 MSAB reports record — no error at the moment of interruption — stated seven years earlier against a different product and build. It is not a finding about current Axiom; it is confirmation that the secondary claim about the 2018 report was accurate.
Reported, but wrong
Omissions are recoverable once you know about them. These are the anomalies where the output looks finished.
A value present in the wrong place, a name whose characters are reversed, a PDF that is listed but will not open — none of these produce an empty cell that prompts a second look.
The clean reports, and how far they reach
Two Graykey reports recorded no anomalies at all.
Scope belongs next to that. The 2025 Graykey and Axiom report covers twelve devices: eleven Apple handsets running iOS 15.1, 15.6 or 17.4.1, and one Samsung Galaxy S22. The 2023 GrayKey Android report covers four devices — Galaxy S22, Galaxy Z Fold3 5G, Galaxy Tab S8 and Pixel 4 — and it did record anomalies, including "Notes were not reported. (Device: Google Pixel 4)" and three social media anomalies.
These are the smallest test sets among the reports we read. A clean report on twelve devices is a real result on twelve devices. It is not evidence about a thirteenth.
Two products we did not find
Cellebrite Premium and Magnet Verakey are marketed on device access and consent-based extraction respectively. We looked for test reports covering them.
We searched the DHS S&T mobile device acquisition index and the NIST CFTT mobile index on 6 September 2026. Neither page's text contains the word "Premium" or the word "Verakey." Neither product appears in any of the fourteen reports we read; Cellebrite is covered through UFED, UFED4PC, Inseyets and Physical Analyzer, and Grayshift through the two Graykey reports.
We are stating an absence from two index pages and fourteen reports on one date. We did not open every report linked from either index, and we make no claim about validation that may exist elsewhere, unpublished, or under another name.
What the dates say, and what they do not
The Cellebrite Inseyets report published on 19 May 2025 carries "September 2024" on its cover. That is roughly an eight-month gap between testing and publication, and it is visible in the hardware: the Apple devices in these test sets run iOS 15.1, 15.6, 16.7, 16.7.3 and 17.4.1, with iOS 18.1 appearing on a single handset in one report.
Counting test-result documents linked from the DHS S&T index on 6 September 2026, by the year of publication:
- 2022 — three
- 2023 — seventeen
- 2024 — two
- 2025 — fourteen
- 2026, through August — one
We are reporting counts, not a trend. CFTT publishes in irregular batches — four reports share a single date in March 2025 — the 2024 figure was followed by fourteen in 2025, and the test specification was updated to version 3.3 in January 2025. No agency we read has stated that testing has slowed or stopped, and we did not find evidence that it has. The counts are here so a reader can see what the published record looks like, and to make one narrower point: independent results describe builds that are, routinely, a year or more behind what is currently sold.
What a CFTT report establishes
Taken together, and stated as carefully as we can:
- A result is version-specific. It describes the build named on the cover, not the product line.
- A result is device-specific. It describes the handsets in the table, at the operating system versions in the table.
- A result is method-specific. The reports say so themselves.
- A result is bounded by the test corpus. An application that was not populated on a device produces "NA," not a finding.
- A clean report does not establish that other versions, devices or methods are clean. An anomalous report does not establish that a product is unreliable in general.
What we could not verify
Not independently verified
Whether any independent validation of Cellebrite Premium or Magnet Verakey exists outside the two index pages and fourteen reports we searched. The 2018 Magnet AXIOM report cited above is a separate historical CFTT result; it is not a Premium or Verakey validation.
Not independently verified
We did not reconcile the report counts published on the NIST CFTT index against those on the DHS S&T index; the two pages link different numbers of documents. Our counts above are from the DHS index only.