NLU-based assistants
This section refers to building NLU-based assistants. If you are working with Conversational AI with Language Models (CALM), this content may not apply to you.Installation RequirementsTo use NLU components, you need to install the For more information about dependency groups, see our Python Versions and Dependencies reference page.
nlu dependency group:Tokenizers
Tokenizers split text into tokens. If you want to split intents into multiple labels, e.g. for predicting multiple intents or for modeling hierarchical intent structure, use the following flags with any tokenizer:intent_tokenization_flagindicates whether to tokenize intent labels or not. Set it toTrue, so that intent labels are tokenized.intent_split_symbolsets the delimiter string to split the intent labels, default is underscore (_).
WhitespaceTokenizer
- Short Tokenizer using whitespaces as a separator
-
Outputs
tokensfor user messages, responses (if present), and intents (if specified) - Requires Nothing
-
Description
Creates a token for every whitespace separated character sequence.
Any character not in:
a-zA-Z0-9_#@&will be substituted with whitespace before splitting on whitespace if the character fulfills any of the following conditions:- the character follows a whitespace:
" !word"→"word" - the character precedes a whitespace:
"word! "→"word" - the character is at the beginning of the string:
"!word"→"word" - the character is at the end of the string:
"word!"→"word"
"wo!rd"→"wo!rd"
a-zA-Z0-9_#@&.~:\/?[]()!$*+,;=-will be substituted with whitespace before splitting on whitespace if the character is not between numbers:"twenty\{one"→"twenty","one"(”{”` is not between numbers)"20\{1"→"20\{1"(”{”` is between numbers)
"name@example.com"→"name@example.com""10,000.1"→"10,000.1""1 - 2"→"1","2"
- the character follows a whitespace:
- Configuration
config.yml
Featurizers
Text featurizers are divided into two different categories: sparse featurizers and dense featurizers. Sparse featurizers are featurizers that return feature vectors with a lot of missing values, e.g. zeros. As those feature vectors would normally take up a lot of memory, we store them as sparse features. Sparse features only store the values that are non zero and their positions in the vector. Thus, we save a lot of memory and are able to train on larger datasets. All featurizers can return two different kind of features: sequence features and sentence features. The sequence features are a matrix of size(number-of-tokens x feature-dimension).
The matrix contains a feature vector for every token in the sequence.
This allows us to train sequence models.
The sentence features are represented by a matrix of size (1 x feature-dimension).
It contains the feature vector for the complete utterance.
The sentence features can be used in any bag-of-words model.
The corresponding classifier can therefore decide what kind of features to use.
Note: The feature-dimension for sequence and sentence features does not have to be the same.
LanguageModelFeaturizer
- Short Creates a vector representation of user message and response (if specified) using a pre-trained language model.
-
Outputs
dense_featuresfor user messages and responses - Type Dense featurizer
- Description Creates features for entity extraction, intent classification, and response selection. Uses a pre-trained language model to compute vector representations of input text.
Please make sure that you use a language model which is pre-trained on the same language corpus as that of your
training data.
-
Configuration
Include a Tokenizer component before this component.
You should specify what language model to load via the parameter
model_name. See the below table for the currently supported language models. The weights to be loaded can be specified by the additional parametermodel_weights. If left empty, it uses the default model weights listed in the table.
- The model architecture is one of the supported language models (check that the
model_typeinconfig.jsonis listed in the table’s columnmodel_name) - The model has pretrained Tensorflow weights (check that the file
tf_model.h5exists, at this time Safetensors are not supported.) - The model uses the default tokenizer (
config.jsonshould not contain a customtokenizer_classsetting)
While the
LaBSE weights are loaded by default for the bert architecture offering a multi-lingual model
trained on 112 languages (see our tutorial and the original
paper), we now recommend using MiniLM model for better performance.The LaBSE weights can still serve as a baseline for initial testing and development. After establishing this
baseline, we strongly encourage exploring optimization with the MiniLM to improve your assistant effectiveness,
before trying to optimize this component with other weights/architectures.rasa/LaBSE weights, which can be found
here:
config.yml
sentence-transformers/all-MiniLM-L6-v2 weights, which can be found
here:
config.yml
RegexFeaturizer
- Short Creates a vector representation of user message using regular expressions.
-
Outputs
sparse_featuresfor user messages andtokens.pattern -
Requires
tokens - Type Sparse featurizer
-
Description
Creates features for entity extraction and intent classification.
During training the
RegexFeaturizercreates a list of regular expressions defined in the training data format. For each regex, a feature will be set marking whether this expression was found in the user message or not. All features will later be fed into an intent classifier / entity extractor to simplify classification (assuming the classifier has learned during the training phase, that this set feature indicates a certain intent / entity). Regex features for entity extraction are currently only supported by the CRFEntityExtractor. -
Configuration
Make the featurizer case insensitive by adding the
case_sensitive: Falseoption, the default beingcase_sensitive: True. To correctly process languages such as Chinese that don’t use whitespace for word separation, the user needs to add theuse_word_boundaries: Falseoption, the default beinguse_word_boundaries: True.
config.yml
CountVectorsFeaturizer
- Short Creates bag-of-words representation of user messages, intents, and responses.
-
Outputs
sparse_featuresfor user messages, intents, and responses -
Requires
tokens - Type Sparse featurizer
- Description Creates features for intent classification and response selection. Creates bag-of-words representation of user message, intent, and response using sklearn’s CountVectorizer. All tokens which consist only of digits (e.g. 123 and 99 but not a123d) will be assigned to the same feature.
-
Configuration
See sklearn’s CountVectorizer docs
for detailed description of the configuration parameters.
This featurizer can be configured to use word or character n-grams, using the
analyzerconfiguration parameter. By defaultanalyzeris set towordso word token counts are used as features. If you want to use character n-grams, setanalyzertocharorchar_wb. The lower and upper boundaries of the n-grams can be configured via the parametersmin_ngramandmax_ngram. By default both of them are set to1. By default the featurizer takes the lemma of a word instead of the word directly if it is available. You can disable this behavior by settinguse_lemmatoFalse.
Option
char_wb creates character n-grams only from text inside word boundaries;
n-grams at the edges of words are padded with space.
This option can be used to create Subword Semantic Hashing.For character n-grams do not forget to increase
min_ngram and max_ngram parameters.
Otherwise the vocabulary will contain only single letters.Enabled only if
analyzer is word.OOV_token.
In this case during prediction all unknown words will be treated as this generic word OOV_token.
For example, one might create separate intent outofscope in the training data containing messages of
different number of OOV_token s and maybe some additional general words.
Then an algorithm will likely classify a message with unknown words as this intent outofscope.
You can either set the OOV_token or a list of words OOV_words:
OOV_tokenset a keyword for unseen words; if training data containsOOV_tokenas words in some messages, during prediction the words that were not seen during training will be substituted with providedOOV_token; ifOOV_token=None(default behavior) words that were not seen during training will be ignored during prediction time;OOV_wordsset a list of words to be treated asOOV_tokenduring training; if a list of words that should be treated as Out-Of-Vocabulary is known, it can be set toOOV_wordsinstead of manually changing it in training data or using custom preprocessor.
This featurizer creates a bag-of-words representation by counting words,
so the number of
OOV_token in the sentence might be important.Providing
OOV_words is optional, training data can contain OOV_token input manually or by custom
additional preprocessor.
Unseen words will be substituted with OOV_token only if this token is present in the training
data or OOV_words list is provided.use_shared_vocab to True. In that case a common vocabulary set between tokens in intents and user messages
is build.
config.yml
sparse_features are of fixed size during
incremental training, the
component should be configured to account for additional vocabulary tokens
that may be added as part of new training examples in the future.
To do so, configure the additional_vocabulary_size parameter while training the base model from scratch:
config.yml
text (user messages), response (bot responses used by ResponseSelector) and
action_text (bot responses not used by ResponseSelector). If you are building a shared
vocabulary (use_shared_vocab=True), you only need to define a value for the text attribute.
If any of the attribute is not configured by the user, the component takes half of the current
vocabulary size as the default value for the attribute’s additional_vocabulary_size.
This number is kept at a minimum of 1000 in order to avoid running out of additional vocabulary
slots too frequently during incremental training. Once the component runs out of additional vocabulary slots,
the new vocabulary tokens are dropped and not considered during featurization. At this point,
it is advisable to retrain a new model from scratch.
The above configuration parameters are the ones you should configure to fit your model to your data.
However, additional parameters exist that can be adapted.
More configurable parameters
More configurable parameters
LexicalSyntacticFeaturizer
- Short Creates lexical and syntactic features for a user message to support entity extraction.
-
Outputs
sparse_featuresfor user messages -
Requires
tokens - Type Sparse featurizer
- Description Creates features for entity extraction. Moves with a sliding window over every token in the user message and creates features according to the configuration (see below). As a default configuration is present, you don’t need to specify a configuration.
- Configuration You can configure what kind of lexical and syntactic features the featurizer should extract. The following features are available:
config.yml
Intent Classifiers
Intent classifiers assign one of the intents defined in the domain file to incoming user messages.LogisticRegressionClassifier
- Short Logistic regression intent classifier, using the scikit-learn implementation.
-
Outputs
intentandintent_ranking -
Requires
Either
sparse_featuresordense_featuresneed to be present. - Output-Example
-
Description
This classifier uses scikit-learn’s logistic regression implementation to perform intent classification.
It’s able to use only sparse features, but will also pick up any dense features that are present. In general,
DIET should yield higher accuracy results, but this classifier should train faster and may be used as
a lightweight benchmark. Our implementation uses the base settings from scikit-learn, with the exception
of the
class_weightparameter where we assume the"balanced"setting. - Configuration
max_iter: Maximum number of iterations taken for the solvers to converge.solver: Solver to be used. For very small datasets you might considerliblinear.tol: Tolerance for stopping criteria of the optimizer.random_state: Used to shuffle the data before training.ranking_length: Number of top intents to report. Set to 0 to report all intents
SklearnIntentClassifier
- Short Sklearn intent classifier
-
Outputs
intentandintent_ranking -
Requires
dense_featuresfor user messages - Output-Example
-
Description
The sklearn intent classifier trains a linear SVM which gets optimized using a grid search. It also provides
rankings of the labels that did not “win”. The
SklearnIntentClassifierneeds to be preceded by a dense featurizer in the pipeline. This dense featurizer creates the features used for the classification. For more information about the algorithm itself, take a look at the GridSearchCV documentation. - Configuration During the training of the SVM a hyperparameter search is run to find the best parameter set. In the configuration you can specify the parameters that will get tried.
config.yml
KeywordIntentClassifier
- Short Simple keyword matching intent classifier, intended for small, short-term projects.
-
Output
s
intent - Requires Nothing
- Output-Example
- Description This classifier works by searching a message for keywords. The matching is case sensitive by default and searches only for exact matches of the keyword-string in the user message. The keywords for an intent are the examples of that intent in the NLU training data. This means the entire example is the keyword, not the individual words in the example.
- Configuration
config.yml
Entity Extractors
Entity extractors extract entities, such as person names or locations, from the user message.If you use multiple entity extractors, we advise that each extractor targets an exclusive
set of entity types. For example, use Duckling to extract dates and times, and
CRFEntityExtractor to extract person names. Otherwise, if multiple extractors
target the same entity types, it is very likely that entities will be extracted multiple times.For example, if you use two or more general purpose extractors like CRFEntityExtractor,
the entity types in your training data will be found and
extracted by all of them. If the slots you are filling with your entity types are of type
text,
then the last extractor in your pipeline will win. If the slot is of type list, then all results
will be added to the list, including duplicates.Another, less obvious case of duplicate/overlapping extraction can happen even if extractors focus on different
entity types. Imagine a food delivery bot and a user message like I would like to order the Monday special.
Hypothetically, if your time extractor’s performance isn’t very good, it might extract Monday here as a time for the order,
and your other extractor might extract Monday special as the meal.CRFEntityExtractor
- Short Conditional random field (CRF) entity extraction
-
Outputs
entities -
Requires
tokensanddense_features(optional) - Output-Example
-
Description
This component implements a conditional random fields (CRF) to do named entity recognition.
CRFs can be thought of as an undirected Markov chain where the time steps are words
and the states are entity classes. Features of the words (capitalization, POS tagging,
etc.) give probabilities to certain entity classes, as are transitions between
neighbouring entity tags: the most likely set of tags is then calculated and returned.
If you want to pass custom features, such as pre-trained word embeddings, to
CRFEntityExtractor, you can add any dense featurizer to the pipeline before theCRFEntityExtractorand subsequently configureCRFEntityExtractorto make use of the dense features by adding"text_dense_feature"to its feature configuration.CRFEntityExtractorautomatically finds the additional dense features and checks if the dense features are an iterable oflen(tokens), where each entry is a vector. A warning will be shown in case the check fails. However,CRFEntityExtractorwill continue to train just without the additional custom features. In case dense features are present,CRFEntityExtractorwill pass the dense features tosklearn_crfsuiteand use them for training. -
Configuration
CRFEntityExtractorhas a list of default features to use. However, you can overwrite the default configuration. The following features are available:
BILOU_flagdetermines whether to use BILOU tagging or not. DefaultTrue.
config.yml
If
pattern features are used, you need to have RegexFeaturizer in your pipeline.If
text_dense_features features are used, you need to have a dense featurizer (e.g. LanguageModelFeaturizer) in
your pipeline.DucklingEntityExtractor
- Short Duckling lets you extract common entities like dates, amounts of money, distances, and others in a number of languages.
-
Outputs
entities - Requires Nothing
- Output-Example
-
Description
To use this component you need to run a duckling server. The easiest
option is to spin up a docker container using
docker run -p 8000:8000 rasa/duckling. Alternatively, you can install duckling directly on your machine and start the server. Duckling allows to recognize dates, numbers, distances and other structured entities and normalizes them. Please be aware that duckling tries to extract as many entity types as possible without providing a ranking. For example, if you specify bothnumberandtimeas dimensions for the duckling component, the component will extract two entities:10as a number andin 10 minutesas a time from the textI will be there in 10 minutes. In such a situation, your application would have to decide which entity type is be the correct one. The extractor will always return 1.0 as a confidence, as it is a rule based system. The list of supported languages can be found in the Duckling GitHub repository. - Configuration Configure which dimensions, i.e. entity types, the duckling component should extract. A full list of available dimensions can be found in the duckling project readme. Leaving the dimensions option unspecified will extract all available dimensions.
config.yml
RegexEntityExtractor
- Short Extracts entities using the lookup tables and/or regexes defined in the training data
-
Outputs
entities - Requires Nothing
- Description This component extract entities using the lookup tables and regexes defined in the training data. The component checks if the user message contains an entry of one of the lookup tables or matches one of the regexes. If a match is found, the value is extracted as entity. This component only uses those regex features that have a name equal to one of the entities defined in the training data. Make sure to annotate at least one example per entity.
When you use this extractor in combination with CRFEntityExtractor, it can
lead to multiple extraction of entities. Especially if many training sentences have entity annotations for
the entity types for which you also have defined regexes. See the big info box at the start of the
entity extractor section for more info on multiple extraction.In the case where you seem to need both this RegexEntityExtractor and another of the aforementioned
statistical extractors, we advise you to consider one of the following two options.Option 1 is advisable when you have exclusive entity types for each type of extractor. To make the
sure the extractors don’t interfere with one another annotate only one example sentence for each
regex/lookup entity type, but not more.Option 2 is useful when you want to use regexes matches as additional signal for your statistical extractor,
but you don’t have separate entity types. In this case you will want to 1) add the
RegexFeaturizer before the extractors in your pipeline 2)
annotate all your entity examples in the training data and 3) remove the RegexEntityExtractor from your pipeline.
This way, your statistical extractors will receive additional signal about the presence of regex matches
and will be able to statistically determine when to rely on these matches and when not to.
-
Configuration
Make the entity extractor case sensitive by adding the
case_sensitive: Trueoption, the default beingcase_sensitive: False. To correctly process languages such as Chinese that don’t use whitespace for word separation, the user needs to add theuse_word_boundaries: Falseoption, the default beinguse_word_boundaries: True.
config.yml
EntitySynonymMapper
- Short Maps synonymous entity values to the same value.
- Outputs Modifies existing entities that previous entity extraction components found.
- Requires An extractor from Entity Extractors
- Description If the training data contains defined synonyms, this component will make sure that detected entity values will be mapped to the same value. For example, if your training data contains the following examples:
New York City and NYC to nyc. The entity
extraction will return nyc even though the message contains NYC. When this component changes an
existing entity, it appends itself to the processor list of this entity.
- Configuration
config.yml
When using the
EntitySynonymMapper as part of an NLU pipeline, it will need to be placed
below any entity extractors in the configuration file.Incremental training
New in 2.2This feature is experimental.
We introduce experimental features to get feedback from our community, so we encourage you to try it out!
However, the functionality might be changed or removed in the future.
If you have feedback (positive or negative) please share it with us on the Rasa Forum.
rasa train --finetune
to initialize the pipeline with an already trained model and further finetune it on the
new training dataset that includes the additional training examples. This will help reduce the
training time of the new model.
By default, the command picks up the latest model in the models/ directory. If you have a specific model
which you want to improve, you may specify the path to this by
running rasa train --finetune <path to model to finetune>. Finetuning a model usually
requires fewer epochs to train machine learning components like DIETClassifier, ResponseSelector and TEDPolicy compared to training from scratch.
Either use a model configuration for finetuning
which defines fewer epochs than before or use the flag
--epoch-fraction. --epoch-fraction will use a fraction of the epochs specified for each machine learning component
in the model configuration file. For example, if DIETClassifier is configured to use 100 epochs,
specifying --epoch-fraction 0.5 will only use 50 epochs for finetuning.
You can also finetune an NLU-only or dialogue management-only model by using
rasa train nlu --finetune and rasa train core --finetune respectively.
To be able to fine tune a model, the following conditions must be met:
- The configuration supplied should be exactly the same as the
configuration used to train the model which is being finetuned.
The only parameter that you can change is
epochsfor the individual machine learning components and policies. - The set of labels(intents, actions, entities and slots) for which the base model is trained should be exactly the same as the ones present in the training data used for finetuning. This means that you cannot add new intent, action, entity or slot labels to your training data during incremental training. You can still add new training examples for each of the existing labels. If you have added/removed labels in the training data, the pipeline needs to be trained from scratch.
- The model to be finetuned is trained with
MINIMUM_COMPATIBLE_VERSIONof the currently installed rasa version.