
In this article, I present my thoughts on the history and future of the internet, centralized and decentralized networks, and ultimatelyâthe possible architecture of a next-generation decentralized network.
There is something wrong with the internet
I first encountered the Internet in 2000. Of course, this is not the very beginningâ the network had already existed before that, but that time can be considered as the first flourishing of the Internet. The World Wide Webâa brilliant invention by Tim Berners-Lee, web 1.0 in its classic canonical form. A multitude of websites and pages, linking to each other with hyperlinks. At first glanceâa simple architecture, like all genius creations: decentralized and free. I can browse through other people's websites, following hyperlinks; I can create my own site, where I publish what interests meâ for example, my articles, photographs, programs, hyperlinks to websites I find interesting. And others link to me.
It seemsâan idyllic picture? But you already know how it all ended.
The number of pages became overwhelming, and finding information turned into a quite non-trivial task. Hyperlinks created by authors simply could not structure this enormous volume of information. Initially, hand-filled directories appeared, followed by massive search engines that began using sophisticated heuristic ranking algorithms. Websites were created and abandoned, information was duplicated and distorted. The internet was rapidly commercializing and moving further away from the ideal academic network. Markup language quickly turned into a formatting language. Advertising emerged, along with annoying intrusive banners and the technology of promotion and tricking search enginesâSEO. The network was quickly cluttered with informational junk. Hyperlinks ceased to be tools of logical connection and transformed into tools of promotion. Websites became encapsulated, shutting themselves off, turning from open
Even then, I had a thought that "something isn't quite right here." A bunch of different sites, ranging from primitive homepages with garish designs to "mega-portals" overloaded with flashing banners. Even if the sites are about the same topic â they are completely unrelated; each has its own design, its own structure, annoying banners, poorly functioning searches, download problems (yes, I wanted to have information offline). Back then, the internet was starting to resemble some form of television, where all sorts of decorations were nailed to useful content.
Decentralization has turned into a nightmare.
What do I want?
Paradoxical as it may seem, even then, without knowing anything about web 2.0 or p2p, I, as a user, didnât need decentralization! Reflecting on my unclouded thoughts of that time, I conclude that I needed... a single database! One where a query would return all results, not just the ones that fit the ranking algorithm the best. One that would have all these results uniformly formatted and styled with my own consistent design, rather than the garish homemade designs of numerous Vasya Pupkins. One that could be saved offline without the fear of the website disappearing tomorrow and losing the information forever. One where I could input my own information â such as comments and tags. One in which I could conduct searches, sorting, and filtering using my personal algorithms.
Web 2.0 and social networks
Meanwhile, the concept of Web 2.0 emerged. Formulated in 2005 by Tim O'Reilly as "a method of designing systems that become better the more people use them through the consideration of network interactions" â it implies active user involvement in the collective creation and editing of online content. Undoubtedly, the pinnacle and triumph of this concept became Social Networks. Massive platforms bringing together billions of users and storing hundreds of petabytes of data.
What have we gained in social networks?
- Interface unification; it turns out that users do not need all the capabilities for creating various eye-catching designs; all user pages have the same design, which everyone finds satisfactory and even convenient; only the content differs.
- Functionality unification; the diversity of scripts also turned out to be unnecessary. 'Feed', friends, albums... over the lifespan of social networks, their functionality has more or less stabilized and is unlikely to change: after all, functionality is determined by the types of activities of people, and people hardly change.
- A unified database; working with such a database turned out to be much more convenient than with numerous fragmented sites; searching has become much simpler. Instead of continually scanning various loosely related pages, caching all of this, ranking according to complex heuristic algorithms â a relatively simple unified query to a single base with a known structure.
- Feedback interface â likes and reposts; in the regular web, Google could not receive feedback from users after they clicked on a link in search results. In social networks, this connection turned out to be simple and natural.
What have we lost? We have lost decentralization, and therefore â freedom.. It is believed that our data no longer belongs to us. Previously, we could host a personal webpage even on our own computer, but now we hand over all our data to internet giants.
Moreover, as the Internet evolved, it attracted the interest of governments and corporations, leading to issues of political censorship and copyright restrictions. Our pages on social networks can be banned and deleted if the content does not comply with certain network rules; an ill-considered post could result in administrative and even criminal liability.
And now we are once again pondering: should we bring back decentralization? But in a different form, one devoid of the shortcomings of the first attempt?
Peer-to-peer networks
The first p2p networks emerged long before web 2.0 and developed in parallel with the evolution of the web. The main classic use of p2p is file sharing; the initial networks were designed for music exchange. Early networks (like Napster) were essentially centralized, which led to them being quickly shut down by rights holders. Followers turned towards decentralization. In 2000, the ED2K protocols (the first eDonkey client) and Gnutella appeared, followed by the FastTrack protocol in 2001 (KaZaA client). Gradually, the degree of decentralization increased, and technologies improved. Torrent systems replaced those with âupload queues,â introducing the concept of distributed hash tables (DHT). As governments tightened regulations, participant anonymity became increasingly important. Development of the Freenet network began in 2000, I2P in 2003, and the RetroShare project launched in 2006. Numerous p2p networks can be mentioned, both those that existed in the past and those currently active: WASTE, MUTE, TurtleF2F, RShare, PerfectDark, ARES, Gnutella2, GNUNet, IPFS, ZeroNet, Tribbler, and many more. There are many of them, and they are quite different â in both purpose and structure⊠Many of you might not even be familiar with all these names. And this is far from everything.
However, p2p networks have many drawbacks. In addition to the technical issues inherent in each specific protocol and client implementation, there is a fairly general drawback â the difficulty of searching (i.e., all the challenges faced by Web 1.0, but in an even more complex variant). There is no Google with its ubiquitous and instant search. While file-sharing networks can still rely on searching by file name or metadata, finding something in overlay networks like onion or i2p is quite challenging, if not impossible.
Overall, if we draw analogies with the classic internet, most decentralized networks are stuck somewhere at the FTP level. Imagine an internet that consists solely of FTP: no modern websites, no web 2.0, no YouTube⊠This is roughly the state of decentralized networks. And despite some isolated attempts to change things, progress has been minimal so far.
Content
Let's turn our attention to another important piece of this puzzleâcontent. Content is the main issue for any online resource, especially decentralized ones. Where can it be sourced from? Of course, one can rely on a handful of enthusiasts (as is the case with existing P2P networks), but then the development of the network will take a considerable amount of time, and the content will be sparse.
Working with the regular internet involves searching for and studying content. Sometimes it means saving it (if the content is interesting and useful, many, especially those who came online during the dial-up eraâincluding myselfâwisely save it offline so it doesnât get lost; after all, the internet is an uncontrollable entity, a website may exist today but be gone tomorrow, a video might be available on YouTube today but deleted tomorrow, and so forth.
For torrents (which we perceive more as a means of delivery than a true P2P network), saving content is an integral part. This is, by the way, one of the problems with torrents: a downloaded file is difficult to transfer to a more convenient location for use (normally, it requires manually regenerating the seed), and it cannot be renamed at all (you can create a hard link, but very few people are aware of this).
In general, many people save content in one way or another. What happens to it afterward? Usually, saved files end up somewhere on the disk, in a folder like Downloads, amidst a chaotic pile, sitting together with many thousands of other files. This is a problemâspecifically for the user. While the internet has search engines, a local user's computer lacks anything similar. Itâs good if a user is organized and accustomed to sorting their âincomingâ downloads. But not everyone is like that...
In reality, there are quite a few people nowadays who donât save anything at all, fully relying on online access. However, P2P networks assume that content is stored locally on the user's device and shared with other participants. Is it possible to find a solution that will engage both categories of users in a decentralized network without changing their habits, and moreoverâmake their lives easier?
The idea is quite simple: what if we create a tool that allows users to save content from the regular internet in a convenient and transparent way, with intelligent saving â using semantic metadata, not just in a general pile, but in a specific structure with the option for further structuring, while simultaneously distributing the saved content in a decentralized network?
Let's start with saving
We won't discuss utilitarian internet usage for checking weather forecasts or flight schedules. We're more interested in self-sufficient and relatively immutable objects â articles (from tweets/posts on social networks to longer pieces like here on Habr), books, images, programs, audio and video recordings. Where does information primarily come from? Usually, it comes from
- social networks (various news, small notes â 'tweets', pictures, audio and video)
- articles on specialized resources (like Habr); there arenât many good resources, and these are usually also structured like social networks
- news websites
Typically, they have standard features: 'like', 'share', 'post to social networks', etc.
Imagine a browser plugin, which will specifically save everything that we liked, reposted, or saved to 'favorites' (or pressed a special button in the plugin's menu, in case that website doesnât have a like/repost/bookmark feature). The main idea is that you just click 'like' â as you have done a million times before, and the system saves the article, picture, or video to a special offline storage, making this article or image available both for offline viewing through the interface of a decentralized client and in the decentralized network itself! I find it very convenient. No extra actions required, and we solve multiple tasks at once:
- saving valuable content that may be lost or deleted
- rapidly populating the decentralized network
- aggregating content from different sources (you could be registered on dozens of internet resources, and all likes/reposts would flow into a single local database)
- structuring interesting content according to your needs rules
It is clear that the browser plugin must be configured for the structure of each site (this is quite feasible â there are already plugins for saving content from YouTube, Twitter, VK, etc.). There are not that many sites for which it makes sense to create personal plugins. Typically, these are popular social networks (and there are barely more than ten) and a few high-quality niche sites like Habr (there are also not many of those). With open code and specifications, developing a new plugin based on a template should not take much time. For other sites, a universal save button could be used, which would save the entire page in mhtml â possibly cleaning the page of ads first.
Now about structuring
By 'smart' saving, I mean at least preserving metadata: content source (URL), a set of previously given likes, tags, comments, their identifiers, etc. Because when saving normally, this information is lost... The source can refer not only to the direct URL but also to the semantic component: for example, a group on a social network or a user who made a repost. The plugin could be smart enough to utilize this information for automatic structuring and tagging. Additionally, it should be understood that the user can always add some metadata to the saved content, for which maximally convenient interface tools should be provided (I have many ideas on how to do this).
Thus, the issue of structuring and organizing the user's local files is resolved. This is already a practical benefit that can be utilized even without any p2p. It is simply an offline database that knows what, where, and in what context we saved, and allows for small research. For example, to find users from an external social network who liked the same posts as you did. Do many social networks allow this explicitly?
It should be mentioned that a single browser plugin is clearly not enough. The second crucial component of the system is the decentralized network service, which operates in the background and serves both the p2p network (network requests and client requests) and the preservation of new content through the plugin. This service, working in tandem with the plugin, will place the content in the correct location, compute hashes (and possibly determine if such content has already been saved previously), and add the necessary metadata to the local database.
Interestingly, the system would already be useful in this form, without any p2p. Many people use web clippers that add interesting content from the web, for example, to Evernote. The proposed architecture is an extended version of such a clipper.
And finally, p2p exchange.
The most pleasant aspect is that information and metadata (both captured from the web and our own) can be exchanged. The concept of a social network fits perfectly into p2p architecture. One could say that social networks and p2p are made for each other. Any decentralized network should ideally be built as a social network; only then will it work effectively. 'Friends' and 'Groups' are essentially the same peers, with which stable connections must be formed, and such connections arise from a natural sourceâshared interests among users.
The principles of preserving and distributing content in a decentralized network are completely identical to the principles of saving (capturing) content from the regular internet. If you are using some content from the network (which means you have saved it), anyone can use your resources (disk space and bandwidth) required to access that specific content.
Likes are the simplest tool for saving and sharing. If I like somethingâwhether in the external internet or within a decentralized networkâthen it means I like the content, and since thatâs the case, I am willing to keep it locally and share it with other participants of the decentralized network.
- Content will not 'disappear'; it is now stored locally with me, and I will be able to return to it later, anytime, without worrying about someone deleting or blocking it.
- I can categorize, tag, comment, and associate it with other content right away or later, essentially creating meaningful metadata â let's call this "metadata formation".
- I can share this metadata with other network participants.
- I can synchronize my metadata with the metadata of other participants.
Perhaps the abandonment of dislikes also seems logical: if I don't like content, it makes sense that I don't want to use my disk space to store it and my internet channel to distribute it. Thus, dislikes don't really fit into decentralization (although sometimes they can be useful). ).
Bookmarks
«» (or "Favorites") â I don't express my attitude towards the content, but I save it in my local bookmarks database. The word "favorites" doesn't quite fit (that's what likes and their subsequent categorization are for), but "bookmarks" fits well. Content in "bookmarks" can also be shared â if it's "needed" (i.e., you use it in some way), it's logical that it may also be "needed" by someone else. Why not use your resources for that?The function of "friends" is quite obvious.
« These are peers, people with similar interests, and thus those who are likely to have interesting content. In a decentralized network, this primarily means subscribing to the news feed from friends and accessing their catalogs (albums) with the content they have saved.Similarly, the "groups" function â collective feeds or forums, or something along those lines, which can also be subscribed to â and thus receive all group materials and share them. Perhaps, like large forums, groups should be hierarchical â this would allow better structuring of group content and limit the flow of information, not accepting or sharing what isn't particularly interesting to you.Everything else
Similarly, the function âgroupsâ refers to collective feeds, forums, or something like that, which can also have subscriptions â allowing users to receive all materials from the group and share them. Perhaps, âgroupsâ, like large forums, should be hierarchical â this would better structure group content, limit the flow of information, and filter out what isn't particularly interesting to you.
Everything else
It should be noted that decentralized architecture is always more complex than centralized ones. In centralized resources, there is a strict dictate of server code. In decentralized ones, there is a need to negotiate among many equal participants. Naturally, this requires the use of cryptography, blockchains, and other advancements primarily developed for cryptocurrencies.
I assume that some form of cryptographic mutual trust ratings may be required, generated by the network participants for each other. The architecture should effectively combat botnets, which, existing in some cloud, can for example self-inflate ratings. It is essential that corporations and botnet farms, despite their technological superiority, do not take control of such a decentralized network; that its main resource consists of real people capable of producing and structuring content that is interesting and useful for other real people.
I would also like such a network to drive civilization toward progress. I have a myriad of ideas on this topic, which do not fit within the scope of this article. I will just say that certain types of scientific, technical, medical content, etc., should take precedence over entertainment content, which will require some form of moderation. Moderating a decentralized network is a non-trivial task but solvable (though the term 'moderation' is entirely inaccurate here and does not reflect the essence of the process â neither externally nor internally... and I haven't even figured out how to name this process).
It would probably be excessive to mention the necessity of ensuring anonymity â both through built-in means (like in i2p or Retroshare) and by routing all traffic through TOR or VPN.
Finally, the software architecture (schematically represented in the image accompanying the article). As mentioned earlier, the first component of the system is a browser plugin that captures content along with its metadata. The second crucial component is a P2P service that operates in the background (the backend). The network's operation should clearly not depend on whether the browser is running. The third component is the client softwareâthe frontend. This can either be a local web service (in which case the user can interact with the decentralized network without leaving their favorite browser) or a standalone GUI application for a specific OS (Windows, Linux, macOS, Android, iOS, etc.). I like the idea of all frontend options coexisting simultaneously. This will also necessitate a stricter backend architecture.
There are many aspects that are not included in this article. Connecting to existing file storage distributions (i.e., when you already have a couple of terabytes downloaded and allow the client to scan it, obtain hashes, match them with what is available in the Network, and join the distribution while also retrieving metadata about your own files from the Networkânormal titles, descriptions, ratings, reviews, etc.), connecting external sources of metadata (such as databases like Libgen), optionally using disk space for storing others' encrypted content (similar to Freenet), the architecture of integration with existing decentralized networks (which is quite complex), the concept of media hashing (using special perceptual hashes for media contentâimages, audio, and videoâwhich will allow matching semantically identical media files that differ in size, resolution, etc.), and much more.
Brief summary of the article
1. In decentralized networks, there is no Google with its search and rankingâbut there is a Community of real people. A social network with its feedback mechanisms (likes, repostsâŠ) and social graph (friends, communitiesâŠ) is the ideal application-level model for a decentralized network.
2. The main idea I bring with this article is the automatic saving of interesting content from the regular internet upon liking/sharing; this can be useful even without p2p, simply for maintaining a personal archive of interesting information.
3. This content can also automatically fill a decentralized network.
4. The principle of automatically saving interesting content also works when liking/sharing within the decentralized network itself.
Source: habr.com
