Originally took from openai/whisperer and rewrote to TS
TypeScript library for normalizing English text. It provides a utility class EnglishTextNormalizer
with methods for normalizing various types of text, such as contractions, abbreviations, and spacing.
EnglishTextNormalizer
consists of other classes you can reuse independently:
EnglishSpellingNormalizer
- uses a dictionary of English words and their American spelling. The dictionary is stored in a JSON file named english.jsonEnglishNumberNormalizer
- works specifically to normalize text from English words to actually numbersBasicTextNormalizer
- provides methods for removing special characters and diacritics from text, as well as splitting words into separate letters.
$ yarn add @shelf/text-normalizer
import {EnglishTextNormalizer} from '@shelf/text-normalizer'
const normalizer = new EnglishTextNormalizer()
console.log(normalizer.normalize("Let's")); // Output: let us
console.log(normalizer.normalize("he's like")); // Output: he is like
console.log(normalizer.normalize("she's been like")); // Output: she has been like
console.log(normalizer.normalize('10km')); // Output: 10 km
console.log(normalizer.normalize('10mm')); // Output: 10 mm
console.log(normalizer.normalize('RC232')); // Output: rc 232
console.log(
normalizer.normalize('Mr. Park visited Assoc. Prof. Kim Jr.')
); // Output: mister park visited associate professor kim junior
$ git checkout master
$ yarn version
$ yarn publish
$ git push origin master --tags
MIT © Shelf