The Debian Project has initiated a general vote on the openness criteria for AI models.

The Debian project has announced a general resolution voting among project developers to establish criteria for accepting machine learning models into the project's main repository. At this stage, a discussion phase has been launched, after which voting will commence (the voting start date has not yet been determined). Approximately a thousand developers involved in package maintenance and the support of Debian infrastructure have the right to vote.

AI models distributed under open licenses, but without providing the source material and tools for training the model, are proposed to be recognized as incompatible with the Debian criteria defining free software (DFSG, Debian Free Software Guideline). If the proposal is approved, such models will not be able to be included in the project’s main repository ('main'). The possibility of supplying these models in the 'non-free' repositories is not considered in the ongoing voting.

Among the issues arising from the lack of data used during training, the following are mentioned:

  • The absence of source data or programs used for training significantly limits the ability to modify finished AI models. Despite the allowance for modification in the license, such modification is practically challenging. Changes may be necessary, for example, when there is a need to replace the tokenizer required for supporting new languages.
  • The data used for training can be viewed as the 'source code' of the model, while the finished model is seen as the output of processing this 'source code' with training tools. Accordingly, for complete model modification, there must be the possibility to modify both the source data and the tools.
  • The inability to reproduce the work performed to create the model without access to the source data and tools.
  • Security and ethical issues. Without source data and tools, the ability to address vulnerabilities in models is limited to binary patches or complete model replacement. Such patches can only be prepared by the model's author, and consumers of the model become entirely dependent on them. Moreover, no one, including the model authors, is capable of understanding the essence of the changes proposed in this way. The absence of source data also complicates the identification of backdoor substitutions in machine learning models.
  • Limitations in study. Without source data, it is impossible to confirm that the model was trained on data provided under licenses that permit such use or to rule out the possibility that data obtained illegally was used during training. Furthermore, if GPL-licensed data was used during training, it may be necessary to analyze whether the outputs generated by the model contain fragments of that data, which require source and license attribution. The developer may inadvertently violate the license of some source data by adding code/content generated by the model to their project.

In October of last year, the OSI (Open Source Initiative) published a definition of an open AI system. An open AI system must provide the following capabilities: use for any purpose without the need for separate permission; examination of the system's operation and inspection of its components; modification for any purposes; transfer to others of both the original and modified versions without limitation on purposes of use. An open AI system must include detailed information about the model architecture, the data used during training, and the training methodology, as well as the source code needed to run and train the AI system. The information must be sufficient for a professional developer to recreate an equivalent AI system independently, using the same or similar data for training.

The Software Freedom Conservancy (SFC), a human rights organization, has criticized such a definition. The dissatisfaction stems from the absence of requirements for providing the data used to train the model. The OSI definition only requires detailed information about the data used for training, not the data itself. The accepted definition guarantees only two of the four proclaimed freedoms of Open Source — the ability to use and distribute, while the abilities to modify and study are not fully ensured.

The OSI's decision is explained by the fact that publishing the original data is often impossible due to reasons beyond the control of the AI model developer, such as the need to maintain confidentiality, the use of copyrighted materials, licensing data from third-party providers, etc. If a requirement to provide data were added, none of the existing large language models would gain open status, and the definition itself would become an unattainable utopia.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster