The Tor network has for years been one of the most recognisable symbols of anonymity on the Internet. Its design, based on multi-layered encryption and passing traffic through a set of independent nodes, is intended to hide the real source of communication. For users this means privacy protection; for court experts and forensic analysts – a challenge that no other protocol poses on such a scale. Analysing data from the Tor network requires not only technical knowledge but also an understanding of what, from an evidentiary standpoint, can be treated as certain, what is merely circumstantial, and what is entirely beyond the reach of identification.
Tor – The Onion Router – works through a system of three or more randomly selected nodes: entry (or guard), relay (intermediate) and exit node. Each of them knows only the neighbouring element of the route – the entry node sees the user's address but does not know the destination, while the exit node knows the address of the server to which the traffic is directed but has no idea who initiated it. This fundamental separation of information means that no single server in the Tor network has a full picture of the data path. In practice this means that if a web server's logs show the IP address of a Tor node, one can state with full conviction that the connection passed through the Tor network, but one cannot indicate who authored it or where the traffic actually came from.
From a forensic-analysis point of view, this distinction is hugely important. Data obtained from system logs – such as the IP address, timestamp, HTTP headers or browser identifier – is purely technical in nature. If the IP address appears in the current or archived Tor consensus (published every hour by the so-called Directory Authorities), the expert can reliably confirm that it is the address of a Tor node, and often also establish whether it acted as a relay or an exit node. It is also possible to determine the period in which it was active, which country it came from, and which operator and autonomous system it belonged to. All of this constitutes hard, verifiable information that can carry high evidentiary value – especially when confirming that particular traffic was anonymised through Tor and therefore did not come directly from the user's device but from an intermediary server.
The problem begins the moment one tries to cross the boundary of that knowledge and identify a specific user. The Tor architecture by design makes such attribution impossible – and it is precisely in this structural feature that its strength lies. Whereas with classic HTTP connections the IP address can be a reliable indicator of a user's identity (or at least of the origin of the connection), with Tor this element loses its evidentiary significance. The network's nodes are most often located in large data centres – at providers such as OVH, Hetzner, M247 or Contabo – and a single node may handle the traffic of hundreds or thousands of users from all over the world. From the destination server's perspective, Tor traffic is therefore impersonal – it merely confirms that someone, using an anonymising network, established a connection to a given resource.
This does not mean, however, that the analysis ends here. There are certain situations in which it is possible to approximately reconstruct the source of the traffic. This most often happens as a result of a user error – for example, opening a link from Tor Browser in an external browser, leaving a trace of their real IP address, or when the system configuration allows a so-called DNS leak or WebRTC leak. It also happens that the event under examination concerns an automated campaign – bots or scripts using Tor to mask traffic – in which case the nature of the repetition, frequency and patterns of the requests makes it possible, with high probability, to assess that an automated process, rather than a manual user, is behind the traffic. All of this, however, requires in-depth heuristic analysis and provides no basis for unambiguously attributing responsibility to a specific person.
In theory, so-called correlation attacks are also possible, consisting of simultaneously monitoring traffic at the entry and exit of the Tor network and comparing the timing and volume characteristics of the packets. In practice, however, these are activities that require control over a significant part of the network infrastructure and enormous analytical resources, available only to specialised agencies. Within the scope of classic forensic or commercial analyses, these methods cannot be applied in any practical way.
That is why, in the expert's assessment – and this is the key conclusion – data from the Tor network has limited but concrete evidentiary value. Its value lies not in allowing the perpetrator to be identified, but in enabling a reliable description of how the communication took place. The IP address of the exit node confirms that the traffic was anonymised; WHOIS data indicates the operator and country; analysis of the Tor consensus makes it possible to establish the period in which the node was active. This information is important in reconstructing events, because it can help distinguish traffic coming from a real user from automated or anonymised traffic. In evidentiary proceedings its significance is comparable to the trace left by a technical mask – one can describe that something was done via a particular tool, but one cannot indicate who held that tool in their hand.
For this reason, when preparing an opinion or technical report, one must very precisely document both the sources of the information and the limitations of the method. It is crucial to keep a copy of the Tor consensus from the time of the event, to describe the method of verification (e.g. queries to the Onionoo API), and to record the date and time in UTC format. It is also worth preserving the original system logs together with their checksums, in order to maintain the full integrity of the material. Only such documentation gives the findings lasting evidentiary value and allows the court or prosecutor to assess the weight and reliability of the opinion.
In summary, data from the Tor network is valuable – but not in the sense in which location data or logs from non-anonymising systems are valuable. It has informational, not identifying, value. It makes it possible to understand the technical context of the event, to indicate that the connection was made through a network of a particular structure and properties, and to show that the IP address cannot be equated with a user. That is enough for it to matter in proceedings – and little enough that it requires caution in interpretation. In an expert's practice, therefore, it constitutes not proof of identity but proof of the method of communication – and as such it should be treated in any proceedings concerning activity carried out through the Tor network.