Vai al contenuto
CercaLettere

La lista di parole e le sue licenze

Questo sito calcola su una lista di parole fissa. Quali parole contiene, da dove vengono e che cosa è successo loro per strada: è tutto qui.

Che cosa contiene la lista

Forme flesse
379.331
Lunghezza
da 2 a 15 lettere
Righe della fonte
505.072
Forme accentate
18.059
Tolte per il contenuto
32
Liste costruite su questo sito
319

Non è una lista ufficiale né una lista da torneo. Contiene forme flesse, termini tecnici e varianti; non contiene tutte le parole dell’italiano, e ciò che non c’è non compare su nessuna pagina di questo sito.

Che cosa è stato cambiato

  • Morph-it! elenca una riga per (forma, lemma, tratti): «porta» compare come nome e come verbo. Uno strumento di parole elenca grafie, quindi le righe diventano una voce per forma.
  • Le voci che non si scrivono interamente con lettere italiane — gruppi con spazio, forme con apostrofo o trattino, sigle con cifre — sono state tolte.
  • 32 forme sono tolte per il contenuto. La regola è aperta, in tools/wordlist/blocklist-it.json, e riguarda solo gli insulti contro caratteristiche protette; il linguaggio volgare resta nella lista, perché sta in ogni dizionario stampato. Tre voci restano schermate con una riserva scritta nel file, perché hanno un omografo innocuo: «giudia» è anche un piatto romano, «checca» anche un sugo, «subnormale» anche un termine di calcolo numerico.
  • L’ordine viene da una lista di frequenza. Decide che cosa viene prima, mai che cosa c’è nella lista.

Perché non ci sono le rime

An Italian rhyme runs from the stressed vowel, Italian orthography writes stress only on a final syllable, and Morph-it! records morphology rather than phonology. The penultimate-stress default would be right most of the time and silently wrong for every sdrucciola, which is worse on a page whose claim is that its rows rhyme.

La strada più ovvia è stata verificata e non porta da nessuna parte: WikiPron ita_latn_broad (CUNY-CL, Apache 2.0 over Wiktionary) è stato scaricato e contiene 89.608 righe, di cui 0 con l’accento tonico segnato. La domanda è reale — 89.560 ricerche al mese — e questa rinuncia è un costo noto, non un’assenza passata inosservata.

Licenze

Le fonti di questa lista hanno licenze aperte, e quelle licenze chiedono che il loro testo accompagni i dati. È quindi riprodotto qui per intero e senza modifiche. Il build di questo sito fallisce se ne manca uno.

The Italian word list on this site — Morph-it! 0.4.8 by Marco Baroni and Eros Zanchetta (SSLMIT, Università di Bologna). This site derives its own list from it: one entry per inflected form, words outside the Italian alphabet dropped, and a content screen applied.

Dual-licensed CC BY-SA 2.0 or LGPL. This site takes the CC BY-SA 2.0 branch. The licensing section of the distribution's own readme is reproduced here verbatim.

LICENSING INFORMATION
======================

This program is dual-licensed free software; you can redistribute it
and/or modify it under the terms of the under the Creative Commons 
Attribution ShareAlike 2.0 License and the GNU Lesser General Public
License.

***********************************************
* Creative Commons Attribution ShareAlike 2.0 *
***********************************************

Morph-it! is licensed under the Creative Commons Attribution
ShareAlike 2.0 License.

You are free:

- to copy, distribute and display the resource;
- to make derivative works;
- to make commercial use of the resource;

under the following conditions:

- you must give the original authors credit;
- if you alter, transform, or build upon this work, you may distribute
  the resulting work only under a license identical to this one;
- for any reuse or distribution, you must make clear to others the
  license terms of this work;
- any of these conditions can be waived if you get permission from the
  copyright holders.

Your fair use and other rights are in no way affected by the above.

You can find a link to the full license from the Morph-it! website.

Copyright (C) 2004-2007 Marco Baroni and Eros Zanchetta.

*************************************
* GNU Lesser General Public License *
*************************************

Morph-it! A free morphological lexicon for the Italian Language
Copyright (C) 2004-2007 Marco Baroni and Eros Zanchetta

This program is free software; you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation; either version 2 of the License, or
(at your option) any later version.

This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the
GNU General Public License for more details.

You should have received a copy of the GNU General Public License along
with this program; if not, write to the Free Software Foundation, Inc.,
51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA.

Estratto da research/9698-word-tools/sources/it/morph-it.tgz → current_version/readme-morph-it.txt · sha256 della fonte b137343dc095e038 · rilevato il 2026-09-02

The word ORDER on every list on this site — FrequencyWords (hermitdave), computed from the OpenSubtitles2018 corpus. It decides which words a reader is shown first and never which words are in the list.

MIT for code, CC-BY-SA-4.0 for content, in the words of the distribution's own README.

# FrequencyWords
Repository for Frequency Word List Generator and processed files

In early days I hosted the generated files on OneDrive with my blog https://invokeit.wordpress.com/frequency-word-lists/ linking to it.
Moving forward, the code and the generated outputs are on GitHub.

### OpenSubtitle tokenized source
The data used to generate 2016 lists can be found at http://opus.lingfil.uu.se/OpenSubtitles2016.php 
The data used to generate 2018 lists can be found at http://opus.nlpl.eu/OpenSubtitles2018.php

### Format
Frequency lists are on the `{word}{space}{numer_of_occurences_in_corpus}`. By example, in file `en_50k.txt` :
```
you 22484400
i 19975318
the 17594291
to 13200962
...
```

### Usages
These data are reused by various widely used opensource projects, among which Wikipedia, input methods and autocomplete keyoards, etc.

### License 
MIT License for code.<br>
CC-by-sa-4.0 for content.

Estratto da research/9698-word-tools/sources/it/frequencywords-README.md · sha256 della fonte a92fcc50e27f7b2d · rilevato il 2026-09-02

Altri dettagli sull’origine dei dati: Fonti.