The following table highlights the benefit of ingredient substitution systems in recipe generation from food image. It compares the performance of ResNet50-based inverse cooking to the Vision Transformer (ViT) and Vision Transformer Multi Layer Sequence (ViT MLS) images encoders. It is based on the substitution system composed of the context encoder, ingredient decoder and ingredient substitution decoder modules presented by the following research contribution: https://orkg.org/paper/R656570/R657576 .
ORKG Comparisons have changed. We have added new features and improved the user interface. Comparisons might look slightly different, but the comparison data itself remains unchanged.