Fetching
arXiv announces new papers once each weekday evening, at 20:00 US Eastern, from Sunday to Thursday. The app fetches once after each announcement and not in between, so refreshing twice does not ask arXiv twice. Requests to the arXiv API are kept at least three seconds apart, as arXiv asks. The other servers have no fixed schedule, so the app asks them for the last few days, through the bioRxiv API, the OSF API and the Crossref REST API.
By default this happens at 05:00, on Wi-Fi while the phone charges, and the digest is built straight after. Opening the app in the morning then only reads a database.
The model
Each paper is represented by the words and two-word phrases in its title and abstract, weighted by TF-IDF. A logistic regression learns from your reactions, with older papers from the phone as negative examples, and is retrained on the phone whenever your reactions change. There is no neural network and nothing to download.
That choice was measured rather than assumed. In a leave-one-out test on a real library of 38 papers, TF-IDF placed 21 of 24 held-out computer vision papers in the top ten of 301 candidates. A small sentence-embedding model, MiniLM, placed 19. At that sample size the difference is not significant, so the simpler model was kept: nothing to download, no native code, and reasons that quote real words.
A later comparison found that combining TF-IDF with a sentence embedding ranks better, by 3 to 12 per cent in nDCG@25, depending on how many papers a reader has liked. It would add 25 to 35 MB and native code to the app, and has not been shipped.
What it learns from
Each thing you do with a paper counts for a fixed amount, from 0 to 1:
| What you did | Counts as |
|---|---|
| Tapped the heart, for more like this | 0.95 |
| Reached the third page of its PDF | 0.9 |
| Shared it | 0.85 |
| Kept its PDF open for 20 seconds | 0.7 |
| Saved it for later | 0.6 |
| Stayed on its abstract for 15 seconds | 0.4 |
| Opened it | 0.25 |
| Tapped the cross, for less like this | 0 |
The strongest thing you did sets a paper's value, and anything else you did moves it part of the way towards 1, never past it. Less like this overrides everything else. The subjects you follow count as short example texts at 0.7.
A paper you scroll past is not used to train the model. It only counts towards how many places its subject gets in later digests.
Putting a digest together
Early on, the model is not trusted much. Its predictions are pulled towards a neutral value in proportion to how little evidence there is: with evidence from 6 papers the model carries about a fifth of the weight, and with 100 about four fifths.
Places in the digest are shared between your subjects by Thompson sampling, from how often you engage with each. Older evidence fades with a half-life of 30 days, so a change of project shows up within weeks. A subject the app knows little about still gets tried, because its uncertainty is wide.
Within a subject, papers are drawn at random with a strong preference for high scores, using Gumbel noise, rather than taken strictly from the top, and a paper too similar to one already chosen is marked down. Refreshing without anything new gives a different digest.
Acceptance at a conference or journal multiplies a paper's score by up to 1.35 at the default setting, so a well-published paper on a topic you dislike still ranks low. The venue comes from what authors write in arXiv's comments and journal fields, or from bioRxiv's and medRxiv's note that a paper has been published. Papers on the Hugging Face daily list get a smaller boost of the same kind.
The first three cards are always the best matches. After them, one card in four is a deliberate detour, each labelled: first a paper from a neighbouring field (math.OC, outside your usual
), then near misses, papers that ranked just below the cut (testing whether this is for you
). How you react to those shows the model where your line is. Digest size, weight on venue, exploration and variety can all be changed in the app's settings.
The reason on each card
For each card, the app finds the paper you kept that it most resembles, among those you liked, saved, downloaded, shared or read, and prints the words the two have in common, strongest first. A card reading matches sparse, autoencoder, diffusion
means exactly that. If nothing you kept has a cosine similarity of at least 0.06, the card names its category instead.
The reason is not taken from the model's weights. Weights fitted on a few dozen papers are unreliable, and printing them produced reasons such as matches optimal, thereby, known
.
Months later
If you passed over a paper three to twelve months ago and it has since been accepted at a conference or journal, it can come back, at most one per digest: You passed on this in March. It was accepted to NeurIPS 2026.
Workshop papers do not count. Once a week, on Wi-Fi, the app asks arXiv for fresh details of those papers.
Citation counts would be the obvious signal, but open databases such as OpenAlex record almost none for preprints this recent, and asking Semantic Scholar or anyone else paper by paper would tell a third party what you read.
Sources and subjects
The app only asks the servers your subjects need, so a mathematician's phone never contacts the Law Archive.
| Server | Fields |
|---|---|
| arXiv | Physics, mathematics, computer science, statistics, engineering, economics, quantitative biology |
| bioRxiv | Biology |
| medRxiv | Medicine and health |
| PsyArXiv | Psychology |
| SocArXiv | Social science |
| EdArXiv | Education |
| Law Archive | Law |
| ChemRxiv | Chemistry, fetched through Crossref |
All 114 subjects, in 14 fields
- Computer science from arXiv
- Machine learning; Language models; Computer vision; Generative models; Robotics; Security and privacy; Systems and networks; Theory and algorithms; Human-computer interaction; Software engineering; Search and recommendation; Databases and data; Graphics and geometry; Networks and society; AI, agents and planning; Neural and evolutionary computing; Hardware and architecture; Information and coding theory; Logic and verification
- Physics from arXiv
- Astrophysics; High energy physics; Condensed matter; Quantum physics; Gravitation and relativity; Fluids and soft matter; Optics and photonics; Chemistry and chemical physics; Biological physics; Plasma and fusion; Earth and atmosphere; Atomic and molecular physics; Nuclear physics; Mathematical physics; Applied and computational physics; Instrumentation and detectors; Physics, society and history; Nonlinear and complex systems
- Mathematics from arXiv
- Analysis and PDEs; Algebra and geometry; Probability; Optimisation; Numerical methods; Combinatorics and number theory; Geometry and topology; Logic and foundations; Functional analysis and operators; Dynamical systems; Groups and representations; Mathematical statistics
- Biology and medicine from arXiv
- Genomics; Neuroscience; Molecular biology; Populations and evolution; Cell and subcellular biology; Tissues, organs and physiology; Molecular networks; Quantitative methods in biology; Medical imaging
- Statistics from arXiv
- Methods and inference; Applied statistics; Computational statistics
- Engineering from arXiv
- Signal processing; Speech and audio; Computational engineering; Control systems
- Biology from bioRxiv
- Cell and molecular biology; Neuroscience and behaviour; Cancer, immunology and disease; Microbiology; Genomics and bioinformatics; Ecology and evolution; Plant biology; Biophysics and bioengineering; Physiology and pharmacology
- Medicine from medRxiv
- Epidemiology and public health; Infectious disease; Neurology and mental health; Cardiovascular and metabolic; Oncology and haematology; Health informatics and imaging; Genetic and genomic medicine; Health systems and policy; Rehabilitation and physiotherapy; Surgery and perioperative care; Women’s and children’s health; Emergency, intensive and palliative care; Other clinical specialties
- Economics and finance from arXiv
- Economics; Finance; Economic theory; Mathematical finance
- Psychology from PsyArXiv
- Cognitive psychology; Clinical and mental health; Social and personality; Developmental psychology; Methods and meta-science
- Social science from SocArXiv
- Sociology; Policy and public affairs; Science and technology studies; Social methods and statistics
- Education from EdArXiv
- Teaching and curriculum; Assessment and education research; Higher and adult education; Subject teaching
- Law from Law Archive
- Public and constitutional law; Criminal and civil law; Business and economic law; International and comparative law; Technology, health and society
- Chemistry from ChemRxiv
- Organic chemistry and synthesis; Physical and computational chemistry; Materials and nanochemistry; Analytical chemistry; Biological and medicinal chemistry
Limits
- It is for Android. There is no iPhone or desktop version.
- It reads preprint servers, not journals, though a card shows when a preprint has since been accepted somewhere.
- Search and library import go through arXiv. Papers from the other servers can be searched once they are on the phone, and a library entry can only be matched to an arXiv paper.
- ChemRxiv papers open in the browser, because ChemRxiv does not let apps download them.
- Papers from PsyArXiv, SocArXiv, EdArXiv and the Law Archive are shown without authors. Asking OSF, which hosts them, for author names made fetching too slow.
- Popular leans heavily towards machine learning, as does the Hugging Face list it ranks by. For many fields it stays empty.
- The first days are general. Until you have reacted to a few papers, the digest leans on your subjects, recency and venue.
How far to trust the numbers
The measurements above come from the developer's own library and from a simulation that runs the app's own ranker for forty mornings with six made-up readers. They are evidence that the approach works, not proof that it works equally well in every field. The full design notes, with every measurement and what it changed, are part of the source code, along with the Python harness that made them.
Thank you to arXiv for use of its open access interoperability.