<?xml version="1.0" encoding="utf-8" ?><rss version="2.0" xml:base="https://rusvectores.org" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>RusVectōrēs</title>
<description>Word embedding models for Russian</description>
<link>https://rusvectores.org</link>
<atom:link rel="self" href="https://rusvectores.org/data/rss.xml/" />
<language>en-EN</language>
<category>Academy</category>
<copyright>Creative Commons Attribution 4.0 International</copyright>
<managingEditor>akutuzov72@gmail.com (Andrey Kutuzov)</managingEditor>
<webMaster>lizaku77@gmail.com (Elizaveta Kuzmenko)</webMaster>
<pubDate>Mon, 18 Jan 2021 13:04:19 +0300</pubDate>
<lastBuildDate>Mon, 31 Mar 2025 11:00:19 +0300</lastBuildDate>
<item>
<title>10 years of RusVectōrēs</title>
<link>https://rusvectores.org</link>
<guid>https://rusvectores.org/ru/anniversary/</guid>
<description>
In April 2025, the RusVectōrēs project turns 10! 
Read our memories and thoughts about this anniversary (in Russian).
</description>
<pubDate>Mon, 31 Mar 2025 11:00:19 +0300</pubDate>
</item>
<item>
<title>RusVectōrēs switched to using a Ukrainian model by default</title>
<link>https://rusvectores.org</link>
<guid>https://rusvectores.org/en/models/#ukrconll_upos_cbow_200_10_2022</guid>
<description>
We are against the war and we stand with Ukraine.
This is why RusVectōrēs is now switched to using the model trained on the Ukrainian Wikipedia and CommonCrawl by default. 
You can of course still choose other models in specific tabs or via API.
</description>
<pubDate>Fri, 29 Jul 2022 17:04:19 +0300</pubDate>
</item>
<item>
<title>New static model trained on RNC and the most recent Russian Wikipedia dump</title>
<link>https://rusvectores.org</link>
<guid>https://rusvectores.org/en/models/#ruwikiruscorpora_upos_cbow_300_10_2021</guid>
<description>
We have published a new static model: ruwikiruscorpora_upos_cbow_300_10_2021.
It is now used as the default query processing model on RusVectōrēs.
ruwikiruscorpora_upos_cbow_300_10_2021 was trained on the Russian National Corpus and Russian Wikipedia dump from November 2021. 
Find all your favorite covid-19 related neologisms!
</description>
<pubDate>Fri, 10 Dec 2021 17:04:19 +0300</pubDate>
</item>
<item>
<title>PCA projections of lexical vectors are now available</title>
<link>https://rusvectores.org</link>
<guid>https://rusvectores.org/visual/</guid>
<description>
The RusVectōrēs visualisation tab now features PCA plots of static lexical vectors (in addition to t-SNE).
Their advantage is them being deterministic: PCA projection for a given list of words and a given model will always be the same.
Also, a bunch of minor bugs is fixed.
</description>
<pubDate>Thu, 26 Aug 2021 17:04:19 +0300</pubDate>
</item>
<item>
<title>Online generation of lexical substitutes from ELMo</title>
<link>https://pypi.org/project/simple-elmo/</link>
<guid>https://rusvectores.org/contextual/</guid>
<description>
We present a new service which generates context-dependent lexical substitutes from pre-trained ELMo models in real time.
You enter a sentence and receive a list of the nearest semantic associates for each word token: something like the same sentence re-worded.
It is important that the associates are context-dependent and will be different for the same word in different sentences. 
This makes it possible to study and demonstrate lexical ambiguity.
Experiment with our contextualized embeddings directly at the RusVectōrēs website!
</description>
<pubDate>Mon, 18 Jan 2021 13:04:19 +0300</pubDate>
</item>
<item>
<title>New fastText and ELMo models; Python library for ELMo</title>
<link>https://rusvectores.org/en/models/</link>
<guid>https://pypi.org/project/simple-elmo/</guid>
<description>New kids in our model repository! 
First, we publish two fastText models trained on the GeoWAC: a large modern web corpus of Russian which is geographically balanced.
The first model is trained on lemmas, the second one is trained on raw tokens: we have never had a static model like this on RusVectōrēs before.
The lemmatized model is also available for queries via our web interface.
Second, we released an ELMo model trained on large Araneum Russicum Maximum corpus. 
Finally, to work with pre-trained ELMo models, you can now use our fully packaged Python library called simple_elmo (https://pypi.org/project/simple-elmo/).
</description>
<pubDate>Thu, 22 Oct 2020 16:04:19 +0300</pubDate>
</item>
<item>
<title>Check out our side projects: RusNLP and ShiftRy</title>
<link>https://nlp.rusvectores.org/</link>
<guid>https://shiftry.rusvectores.org/</guid>
<description>In the first day of autumn we invite you to check out our side projects.
First, RusNLP (https://nlp.rusvectores.org/) is a search engine for papers presented in Russian NLP conferences.
Second, ShiftRy (https://shiftry.rusvectores.org/) is a web service for analyzing diachronic changes in the usage of words in news texts from Russian mass media.
</description>
<pubDate>Tue, 01 Sep 2020 16:03:19 +0300</pubDate>
</item>
<item>
<title>Large ELMo model trained on the Tayga corpus, and the accompanying code.</title>
<link>https://rusvectores.org/</link>
<guid>https://github.com/ltgoslo/simple_elmo</guid>
<description>Large (LSTM size 2048x2) ELMo model trained on the Tayga corpus is published, along with the code to work with such models.
The archive with the model now contains raw TensorFlow checkpoints, in case you need them. Also, ELMo models are now additionally evaluated with the paraphrase detection task.
</description>
<pubDate>Fri, 31 Jan 2020 05:03:19 +0300</pubDate>
</item>
<item>
<title>Dynamic interactive graphs are added to the web interface</title>
<link>https://rusvectores.org/</link>
<guid>https://github.com/lizaku/vec2graph</guid>
<description>Along with the word nearest neighbors' lists, we now feature dynamic interactive graphs showing relations between these neighbors.
They can be tuned with regards to the edge building threshold, allowing experiments with ambiguous words.
To build graphs, the vec2graph library is used.
</description>
<pubDate>Fri, 22 Nov 2019 09:03:19 +0300</pubDate>
</item>
<item>
<title>ELMo models are published</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/models/</guid>
<description>New contextualized ELMo models for Russian are published: trained on tokens or lemmas, for you to choose.
Put your words vectors in context! See more on ELMo here: https://allennlp.org/elmo
</description>
<pubDate>Mon, 26 Aug 2019 09:03:19 +0300</pubDate>
</item>
<item>
<title>Low-frequency associates are now hidden by default</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/associates/</guid>
<description>Nearest associates lists in all tabs now default to showing only high-frequency and mid-frequency associates. 
There are frequency checkboxes to change this behaviour: for example, one can turn on showing low-frequency words. 
This makes it easy to fine-tune the balance between coverage and quality.
</description>
<pubDate>Mon, 22 Apr 2019 11:03:19 +0300</pubDate>
</item>
<item>
<title>RusVectōrēs tutorial notebook is updated.</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/rusvectores5/</guid>
<description>RusVectōrēs tutorial notebook is updated. It shows how to preprocess Russian words in order to look them up in our models, and how to work with the models themselves. It takes into account the changes introduced in 2019.
https://github.com/akutuzov/webvectors/blob/master/preprocessing/rusvectores_tutorial.ipynb
</description>
<pubDate>Mon, 28 Jan 2019 11:03:19 +0300</pubDate>
</item>
<item>
<title>New set of models is released, many significant changes</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/rusvectores5/</guid>
<description>In January 2019, we released a set on new word2vec and fastText models for Russian. You already can download them, and they are used in our web interface. 
We also switched to new model delivery schema (thanks for the support from the NLPL), increasing speed and realiability of your downloads. </description>
<pubDate>Fri, 18 Jan 2019 11:03:19 +0300</pubDate>
</item>
<item>
<title>Summary of RusVectōrēs users survey published</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/survey/</guid>
<description>We published the summary of RusVectōrēs users survey, look at it! We plan to start implementing the new visualization features according to the demand shown in the survey.</description>
<pubDate>Fri, 21 Dec 2018 01:03:19 +0300</pubDate>
</item>
<item>
<title>User survey about RusVectōrēs</title>
<link>https://rusvectores.org/</link>
<guid>https://docs.google.com/forms/d/1tZrrL7Va0PqfJBmwA4Icmp5gPVnE-vgPNxm5MxVFZ1A</guid>
<description>
Your feedback is important! With the help from HSE master students, we ask you to kindly answer a few questions about RusVectōrēs (in Russian).
This survey will help us to improve the service, particularly with regards to visualizations.
https://docs.google.com/forms/d/1tZrrL7Va0PqfJBmwA4Icmp5gPVnE-vgPNxm5MxVFZ1A
</description>
<pubDate>Tue, 27 Nov 2018 01:03:19 +0300</pubDate>
</item>
<item>
<title>RusVectōrēs bot in Telegram</title>
<link>https://rusvectores.org/</link>
<guid>http://t.me/rusvectores_bot</guid>
<description>We launched the RusVectōrēs bot in the Telegram messenger service. 
You can ask question to this bot while commuting to your office/university, and it will send queries to our API.
This can be convenient in the situations when you want to check upon an idea, but no laptop is available nearby.
http://t.me/rusvectores_bot
</description>
<pubDate>Sat, 22 Sep 2018 01:03:19 +0300</pubDate>
</item>
<item>
 <title>Long-expected model trained on the Taiga corpus is published</title>
<link>https://rusvectores.org/</link>
<guid>https://tatianashavrina.github.io/taiga_site/</guid>
<description>Taiga is a large structured web corpus of Russian aiming to provide a well-maintained and exhaustive resource for training machine learning models (https://tatianashavrina.github.io/taiga_site/).
We present a Continuous Skipgram word embedding model trained on the full Taiga (excluding the poetry subcorpus): taiga_upos_skipgram_300_2_2018.</description>
<pubDate>Wed, 20 Jun 2018 01:03:19 +0300</pubDate>
</item>
<item>
 <title>Tutorial on text preprocessing for RusVectōrēs models released</title>
<link>https://rusvectores.org/</link>
<guid>https://github.com/akutuzov/webvectors/blob/master/preprocessing/rusvectores_tutorial.ipynb</guid>
<description>We have developed a tutorial (https://github.com/akutuzov/webvectors/blob/master/preprocessing/rusvectores_tutorial.ipynb) explaining how to convert raw Russian text to tagged lemmas suitable for RusVectōrēs models, how to perform basic operations on word embeddings and how to use the RusVectōrēs API.</description>
<pubDate>Sat, 12 May 2018 01:03:19 +0300</pubDate>
</item>
<item>
 <title>Our models score high in the RUSSE'18 shared task; new fastText model released</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/models/#araneum_none_fasttextcbow_300_5_2018</guid>
<description>RusVectōrēs models won top rankings in the RUSSE'18 word sense induction shared task (https://arxiv.org/abs/1803.05795v1). Additionally, check the updated fastText model trained on the Araneum corpus; now it uses not only 3-grams, but also 4-grams and 5-grams.</description>
<pubDate>Mon, 26 Mar 2018 02:03:19 +0300</pubDate>
</item>
<item>
 <title>Review of RusVectōrēs new features in 2017.</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/rusvectores4/</guid>
<description>New fastText models, word frequencies indication, and proper names in the PoS tags: learn about the new features introduced to RusVectōrēs in the Fall of 2017.</description>
<pubDate>Fri, 05 Jan 2018 02:03:19 +0300</pubDate>
</item>
<item>
 <title>Check the model trained on Araneum Russicum web corpus!</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/models/#araneum</guid>
<description>We added the model trained on Araneum Russicum Maximum, which is one of the largest Russian web corpora (more than 10 billion words). In addition, all the models are re-evaluated using RuSimLex965 semantic similarity test set (it is more consistent than RuSimLex999).</description>
<pubDate>Wed, 09 Aug 2017 14:03:19 +0300</pubDate>
</item>
<item>
 <title>Visualizations are substantially upgraded</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/visual</guid>
<description>Visualizations are substantially upgraded. You can now use several word sets as an input: they will be labeled with different colors in the visualizations. If there is only one set, the colors will correspond to the words' parts of speech. Additionally, it is now possible to visualize your data in TensorFlow Embedding Projector with one mouse click.</description>
<pubDate>Fri, 30 Jun 2017 14:03:19 +0300</pubDate>
</item>
<item>
 <title>It is now possible to download archived models</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/models</guid>
<description>We present a separate page for models where you can download both recent and archived models, and compare them against each other. Also, we added the links to Russian test sets and to the conversion table from Mystem to the Universal PoS Tags.</description>
<pubDate>Thu, 09 Mar 2017 20:03:19 +0300</pubDate>
</item>
<item>
 <title>Screencast about RusVectōrēs and review of new features</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/rusvectores3</guid>
<description>Learn about the new features we introduced in 2016 and watch the new screencast video on working with RusVectōrēs.</description>
<pubDate>Sun, 12 Feb 2017 20:03:19 +0300</pubDate>
</item>
<item>
 <title>New models and new PoS tags</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/about#models</guid>
<description>Major update of the models: news corpus now covers events up until November 2016, Wikipedia dump is updated to the same date. Additionnally, PoS tags are converted to the Universal Parts of Speech standard, and the models' vocabularies now feature multi-word-entities (bigrams).</description>
<pubDate>Thu, 02 Feb 2017 20:03:19 +0300</pubDate>
</item>
<item>
 <title>Semantic similarity API</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/about</guid>
<description>API now allows to get semantic similarity of word pairs. Query format: https://rusvectores.org/MODEL/WORD1__WORD2/api/similarity/</description>
<pubDate>Fri, 18 Nov 2016 03:03:19 +0300</pubDate>
</item>
<item>
 <title>Hints as you type a query</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/similar</guid>
<description>There are now hints as you type a query. Note that there are words not covered by hints as well (though they are rare and strange)!</description>
<pubDate>Sat, 22 Oct 2016 03:03:19 +0300</pubDate>
</item>
<item>
 <title>Training models on user-supplied corpora is now disabled</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/contacts</guid>
<description>For security reasons, the option to automatically train models on user-supplied corpora is now disabled. However, if you have an interesting corpus, contact us, and we will be glad to train a model for you.</description>
<pubDate>Fri, 01 Jul 2016 16:03:19 +0300</pubDate>
</item>
<item>
 <title>RusVectōrēs source code is now open and available.</title>
<link>https://rusvectores.org/</link>
<guid>https://github.com/akutuzov/webvectors</guid>
<description>RusVectōrēs source code is released on Github as Webvectors framework (https://github.com/akutuzov/webvectors). We welcome comments, bug reports and pull requests.</description>
<pubDate>Thu, 07 Apr 2016 06:03:19 +0300</pubDate>
</item>
<item>
 <title>API now provides output in JSON.</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/news/%D1%83%D0%B4%D0%B0%D1%80/api/json</guid>
<description>We provide a simple API to get the list of semantic associate for a given word in a given model. There are two possible formats: json and csv. Perform GET requests to URLs following the pattern https://rusvectores.org/MODEL/WORD/api/FORMAT where MODEL is the identifier for the chosen model, WORD is the query word and FORMAT is "csv" or "json", depending on the output format you need. We will return a json file or a tab-separated text file with the first 10 associates.</description>
<pubDate>Mon, 04 Apr 2016 06:03:19 +0300</pubDate>
</item>
<item>
 <title>Service for Englisn and Norwegian is launched</title>
<link>https://rusvectores.org/</link>
<guid>http://ltr.uio.no/semvec</guid>
<description>Web service with distributional models for English and Norwegian is launched, based on RusVectōrēs engine.</description>
<pubDate>Tue, 15 Mar 2016 16:03:19 +0300</pubDate>
</item>
<item>
 <title>Fixed a bug in user model training</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/upload</guid>
<description>We fixed a bug because of which training user models was broken.</description>
<pubDate>Wed, 03 Feb 2016 17:48:19 +0300</pubDate>
</item>
<item>
 <title>RusVectōrēs 2.0: Christmas Edition</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/en/christmas</guid>
<description>RusVectōrēs 2.0: Christmas Edition is officially released.</description>
<pubDate>Wed, 23 Dec 2015 00:37:19 +0300</pubDate>
</item>
<item>
 <title>News model is updated.</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/about#models</guid>
<description>News model is now trained on texts up to November 2015</description>
<pubDate>Thu, 17 Dec 2015 15:37:19 +0300</pubDate>
</item>
<item>
  <title>Query PoS filtering</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/similar</guid>
<description>In Similar Words one can now filter results with the query part of speech.</description>
<pubDate>Tue, 15 Dec 2015 21:20:58 +0300</pubDate>
</item>
<item>
  <title>API released</title>
<link>https://rusvectores.org/</link>
<guid>https://rusvectores.org/news/%D1%83%D0%B4%D0%B0%D1%80/api</guid>
<description>API is implemented. It outputs 10 nearest neighbors for given word and model. Example: https://rusvectores.org/news/удар/api</description>
<pubDate>Fri, 11 Dec 2015 21:20:58 +0300</pubDate>
</item>
</channel>
</rss>