Kees Cook, fost administrator principal al sistemului la kernel.org și lider al echipei de securitate Ubuntu, a demonstrat posibilitatea creării unui commit al cărui ID scurt coincide cu un commit anterior adăugat în nucleul Linux. Experimentul a fost realizat ca o dovadă a fezabilității tranziției la ID-uri scurte de commit de 16 caractere în nucleul Linux, despre care s-a discutat anterior pe lista de discutii a dezvoltatorilor nucleului, dar care nu a fost aprobat de Linus Torvalds.
ID-urile scurte ale commit-urilor sunt generate prin păstrarea primelor 12 caractere din hash-ul SHA-1 (48 de biți din 160 de biți). Având în vedere că numărul de obiecte în nucleu identificate prin hash-ul SHA-1 a depășit 13 milioane, apariția coliziunilor utilizând un prefix de 12 caractere a devenit o chestiune de timp. Ca exemplu, sunt prezentate deja obiecte adăugate în nucleu care se suprapun prin ID-urile lor de 11 caractere. În plus, s-a menționat că suprapunerea ID-urilor de 12 caractere a fost deja înregistrată în luna octombrie, dar înainte de trimiterea patch-ului, utilitarul checkpatch a descoperit problema.
ID-urile scurte sunt utilizate la publicarea de linkuri scurte către commit-uri și sunt specificate la trimiterea modificărilor în eticheta „Fixes”, ca referință la commit-ul în care problema a fost rezolvată în patch-ul trimis (de exemplu, „Fixes: e21d2170f366”). Apariția coliziunilor, în care mai multe modificări diferite sunt asociate cu același ID scurt, poate duce la problematizarea instrumentelor pentru analiza și validarea modificărilor care țin cont de conținutul etichetelor „Fixes”. De exemplu, aceste etichete sunt considerate în procesorul check_fixes, utilizat în ramura linux-next, precum și în scripturile de analiză a corectării vulnerabilităților și urmărirea ciclului de viață al patch-urilor.
Linus Torvalds has expressed skepticism about the proposal to increase the minimum size of shortened identifiers, as the actual number of commits in the repository is significantly lower than the number of objects (approximately 1/8). Most likely, if random collisions occur, they will be between a commit and an object of a different type (for example, a blob or a branch). In his opinion, shortened identifiers are meant to be clear, readable, and easily quotable, and there are currently no objective reasons to increase their size.
One of the developers suggested achieving a reduction in size while increasing the number of significant bits by using a new format based on Base36 encoding (characters 0-9a-z) instead of hexadecimal digits. According to Linus, such a change would create more problems than it solves. For example, support for the new format would need to be added to existing utilities, and a format identifier would need to be introduced to distinguish between the old and new formats.
To demonstrate that the issue with shortened identifiers is not theoretical and its resolution should not be delayed, Kees Cook prepared a documentation change for the kernel, the shortened identifier of which (1da177e4c3f4) matched the commit identifier for creating the kernel branch 2.6.12-rc2. The collision was found after 6 hours of calculations on a system with an NVIDIA GeForce RTX 3080 GPU.
The search was conducted using the lucky-commit tool — random spaces were added to the target patch text until the 12-character SHA-1 prefix matched the existing prefixes of commits in the kernel. According to Kees, the problem lies not so much in random collisions but in the potential manipulation of shortened identifiers for malicious purposes, such as bypassing certain checks.
Sursa: opennet.ro
