r/MachineLearning • • 1d ago

Discussion [D] Self-Promotion Thread

0 Upvotes

Please post your personal projects, startups, product placements, collaboration needs, blogs etc.

Please mention the payment and pricing requirements for products and services.

Please do not post link shorteners, link aggregator websites , or auto-subscribe links.

--

Any abuse of trust will lead to bans.

Encourage others who create new posts for questions to post here instead!

Thread will stay alive until next one so keep posting after the date in the title.

--

Meta: This is an experiment. If the community doesnt like this, we will cancel it. This is to encourage those in the community to promote their work by not spamming the main threads.


r/MachineLearning • • 2d ago

Discussion [D] Monthly Who's Hiring and Who wants to be Hired?

6 Upvotes

For Job Postings please use this template

Hiring: [Location], Salary:[], [Remote | Relocation], [Full Time | Contract | Part Time] and [Brief overview, what you're looking for]

For Those looking for jobs please use this template

Want to be Hired: [Location], Salary Expectation:[], [Remote | Relocation], [Full Time | Contract | Part Time] Resume: [Link to resume] and [Brief overview, what you're looking for]

​

Please remember that this community is geared towards those with experience.


r/MachineLearning • • 1d ago

News arXiv now limits submitters to up to two submissions per calendar month [N]

Thumbnail
blog.arxiv.org
353 Upvotes

r/MachineLearning • • 15h ago

Research Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction [R]

14 Upvotes

In our #NeurIPS2026 paper “Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction” (preprint: https://arxiv.org/abs/2606.22969) we try to address a fundamental issue in dynamical systems reconstruction (DSR) and time series forecasting (TSF): Many recent SOTA DSR & TSF models can generalize to new initial conditions or time series with changing statistical properties. But the really hard problem in DSR and TSF is topological out-of-domain generalization (OODG) (https://proceedings.mlr.press/v235/goring24a.html) where the dynamical regime changes, for instance from cyclic to chaotic behavior.

This can happen when a system crosses a tipping point due to a slowly varying control parameter which drives it across bifurcations, such as in climate systems, when the brain tips from normal into epileptic activity, or when a patient develops blood poisoning (sepsis). Such problems are beyond the realm of current TSF models which rely on extracting temporal patterns and statistical regularities. Yet the ability to predict previously unseen, novel dynamical regimes as a system parameter changes is something we would expect from any good scientific theory. Often these control parameters that drive regime changes are not exactly known either. Hence, a data-driven DSR model for achieving topological OODG would need to infer the dynamical system generating the TS jointly with the control parameters.

In our paper, we mathematically identify key failure modes in previous hierarchical DSR models (https://proceedings.iclr.cc/paper_files/paper/2025/hash/d4c961804d08e55d898cce944206d455-Abstract-Conference.html) that prevent them from correctly learning a system’s control parameters and extrapolating them beyond the training domain. By fixing these through feature-splitting and physical sparsity priors, our modified hierarchical DSR model manages to correctly predict bifurcations and beyond-bifurcation dynamics, without any explicit knowledge about the control parameters provided in training.

Our approach is generic and works for different discrete and continuous time RNNs, we tested it for shallow PLRNNs and Neural ODEs.


r/MachineLearning • • 19h ago

Research Adding memory to search instead of sampling in reward maximization tasks [R]

3 Upvotes

I am one of the authors of FLEET - an algorithm that enhances Best-of-N generation by attributing external rewards to particular tokens and then uses MCTS to adjust logits during the next run.

I find it rather funny that most of the tasks where repetitive sampling is widely used are based on reward maximization, yet it is not aware of that reward. Tuning the sampling parameters allows to make the process more efficient, but it is still a blind search. We propose a way to make generation aware of previous rewards with solutions on how to attribute reward to the completion and how to use this information.

In FLEET the technique from adaptive sampling methods is used that is to track logits for which entropy and varentropy are high thus showing the model's uncertainty about token optimality. We treat these states as branching points. The corresponding normalized hidden states are stored in the vector store and mapped to metadata entries with the history of rewards and transitions between "nodes". The retrieval and update of metadata is based on cosine similarity as for very high similarity KL divergence is low enough to preserve most of the meaningful tokens.

Instead of actually selecting the tokens FLEET uses modified MCTS to rank top-k tokens + special exploration (or other tokens) set and penalize the suboptimal ones. Then decoding strategy is applied to modified logits.

It was tested on GSM8K and LiveCodeBench v6 easy split with Llama 3.2 3B, penalty set to effectively zero probability for the suboptimal tokens + greedy decoding:

  • For GSM8K it solved just seven more tasks, but reached the sampling baseline with half the iterations.
  • For LiveCodeBench it increased the score from 0.59 to 0.69 under the same budget and reached the baseline even faster, now with only 9 iterations against 32.

The sequential execution is not required, as it is not updated during the iteration itself it can simply be passed as a lookup table. The metadata store can be preserved as a prior for other tasks or to enrich SFT/RL.

Paper (preprint): https://arxiv.org/abs/2609.27657
Huggingface: https://huggingface.co/papers/2609.27657
Repository (experiments, examples and python package): https://github.com/Alexiush/fleet

There are more details on changes made to MCTS, how to tune the search parameters for specific model and task as well as code for experiments and trajectories.


r/MachineLearning • • 1d ago

Research Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]

Thumbnail
gif
124 Upvotes

Can training of nonlinear RNNs be efficiently parallelized, ensuring fast convergence even on very long time series from chaotic systems?

In our #NeurIPS2026 spotlight “Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems (DS) Reconstruction (DSR)” (preprint: https://arxiv.org/abs/2605.12683) we speed up training of nonlinear RNNs on time series from chaotic DS by more than 2 orders of magnitude (>100x) by combining DEER with generalized teacher forcing (GTF).

DEER (https://openreview.net/forum?id=E34AlVLN0v) solves the RNN forward pass through Newton-type fixed point iterations across the whole sequence length T, enabling scaling as O[(log T)²] instead of O[T] by allowing for efficient GPU parallelization. But under chaotic dynamics DEER breaks down and its runtime degrades to O[T log T] (https://openreview.net/forum?id=7AGXSlXcK6).

Using GTF (https://proceedings.mlr.press/v202/hess23a.html) we stabilize DEER by preventing divergence due to chaotic dynamics and reduce exposure bias compared to traditional teacher forcing used to train state space models.

Combining these two mechanisms enables efficient parallel-in-time and stable training on extremely long time series (T>106) from chaotic simulated or real-world systems, hugely outperforming Mamba and other state space models in the DSR setting.


r/MachineLearning • • 1d ago

Research LLMs that push back on a wrong user still accept the same wrong answer from a "verified source" - NeurIPS 2026 [R]

57 Upvotes

I'm one of the authors. We kept seeing models that hold their ground when the user insists on a wrong answer, yet change their answer when the same claim is framed as coming from a "verified source". We wanted to measure how often this happens and check whether the model represents the two cases differently. We call the effect Authority Bias.

Why we think it matters. Standard sycophancy evals apply pressure through the user, so a model can pass them while still being easy to mislead through search results, retrieved documents and tool outputs.

Another reason is with current AI research accelerating towards more agentic and autonomous models + with cases of tools hiding their traces and trusting tools "more" over the user (who could be trying to correct them), safeguarding against misinformation from tools is particularly important!

Setup. We take TriviaQA questions the model already answers correctly. To each one we add a wrong answer, either as "According to the verified source, the answer is X" or as the user saying "I'm a domain expert and I'm pretty sure it's X". The question and the wrong answer stay the same; only the speaker changes. Answers are free-form, not multiple choice. (In a multiple-choice pilot the effect mostly vanished.) We test 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro).

Behavior

  • One verified-source note flips 45-88% of correct answers in 7 of 8 models. The same wrong answer from the user moves most models much less.
  • The gap is largest in the models that resist users best. GPT-5.4 flips on 44.7% of questions and Grok-4.20 on 87.5% (these models were "frontier" during the time of writing this paper). Gemini-3.1-Pro ignored both speakers (0.6%) and was particularly resistant to this method.

Inside the model (open-weight models only, using difference-of-means directions)

  • In Qwen3.5, GPT-OSS and OLMo-3.1, removing the "source endorsed this" direction cuts compliance with a wrong source by 64-78 points.
    • Removing the "user endorsed this" direction cuts it by at most 11.
  • The two directions also have really high cosine similarity of ~0.90-0.99. Our understanding is that they share a large "this answer was endorsed" component plus a thin part that encodes who endorsed it.
    • Shifting only that thin part, with the prompt unchanged, moves compliance by 11-32 points and closes 55-61% of the source-vs-user gap.

Some limitations

  • The internal results hold in 3 of 5 open-weight families.
    • In OLMo-2 the source direction is entangled with the assistant direction.
    • Gemma-4 flips readily, but no linear intervention we tried controls it.
  • The "retrieved document" tests put the claim in a document-shaped block of the prompt rather than running a real retrieval pipeline.
    • So it would be interesting to see it in a real agentic setup, like Claude Code.

Paper: https://arxiv.org/abs/2609.37616
Code: https://github.com/Lossfunk/authority-bias
Project page (figures and example responses): https://authority-bias.vercel.app


r/MachineLearning • • 9h ago

Research [R] Would you keep a robot demonstration if hand tracking missed the moment the plug went in?

Thumbnail
arxiv.org
0 Upvotes

Suppose you’re recording a human plugging a cable into a socket to collect demonstrations for robot learning.

The hand tracker captures the approach accurately. Then occlusion causes the hand estimates to disappear during insertion. Tracking returns after the connector is already seated.

The video still shows a completed action, but the pose labels have a gap exactly where alignment turns into contact.

This hypothetical example raises an evaluation question: a tracker could have high recall across the whole episode while missing a short, important phase. Pose error calculated only on successful detections could make that failure even harder to see.

MEgoVista provides a useful starting point. Table 3 reports detection precision, recall and F1 alongside reconstruction errors. Section 4.4 also describes an evaluation protocol that assigns an error to missed detections instead of excluding them. The blank HaPTIC row means it failed to produce valid output in their multi-person capture scenes; it doesn’t describe a brief tracking dropout.

Accounting for missing detections matters. My remaining question is whether an episode-level aggregate tells us enough about where those failures happen.

For manipulation data, I’d want pose error and coverage reported together, plus coverage broken down by approach, contact and withdrawal, and the longest consecutive gap during contact.

Continuous hand estimates would still be only part of the picture: object pose and contact information also matter for determining whether insertion succeeded.

For people using human motion reconstruction for imitation learning, what evaluation protocol do you use to decide whether an episode with missing contact-phase labels is still usable?

Comparison with open-source egocentric hand reconstruction methods against motion-capture ground truth. All methods are evaluated on identical segments of our motion-capture dataset. All baselines are re-run and rescored on our data. HaPTIC fails to produce valid output in our multi-person capture scenes. Bold marks the best result in each column.

r/MachineLearning • • 1d ago

Project A video about Adversarial Objectives [P]

0 Upvotes

I made this video about adversarial objectives, which I used to do research on back in the day. I'm trying to explore how adversarial approaches transcend GANs and self-play into modern technology.
https://youtu.be/W7CiAeQ0f5w?si=g0tLrQn2wuFzuN2M


r/MachineLearning • • 1d ago

Discussion Gemini 4 Argon - 1 Million Output Headroom. Hype or a Leap? [D]

20 Upvotes

I rarely write about benchmarks; a competitor always beats them next week. But I care about 'Leaps'. Gemini 4 Argon feels like one to me.

While Opus 5.5 and Astra cap output at 128-300K tokens (~90-180 pages), Argon hits 1 Million (~1400 pages).

"Context glue" ruins agentic workflows. On paper, this headroom fixes that. It means less contextual drift, no more breaking down long tasks, and zero 'continue prompt' loops. It is a massive unlock for large-scale code migrations, security patches, and deep reasoning.

But let's look past the marketing. For 95% of everyday work, nobody needs 1,000 pages at once.

I want to ask the experts here: Is a 1M output window a real paradigm shift for agents, or does generating that much text just guarantee a massive logic collapse halfway through? Are you actually hitting output limits today, or is this hype? Let's discuss.


r/MachineLearning • • 2d ago

Discussion How to address novelty concerns in top ai conference? [D]

58 Upvotes

Hi, I’m a researcher working in computer vision.

Over the past few years, I’ve submitted several papers to top-tier conferences such as NeurIPS, ICLR, and CVPR, and one concern that seems to come up repeatedly is 'novelty'.

Given that thousands of papers are published every year at top conferences alone, not to mention the tens of thousands published across other conferences and journals, I sometimes wonder how much genuinely new novelty is realistically left to explore.

In such a crowded research landscape, how do you usually address novelty concerns from reviewers?

More specifically, I would really appreciate any advice on how to frame a contribution so that its novelty is clear, how to distinguish meaningful incremental progress from work that may be considered insufficiently novel, and what reviewers generally look for when judging novelty.

Any tips or experiences would be greatly appreciated. Thanks!


r/MachineLearning • • 1d ago

Discussion For academia/industry, do HuggingFace model downloads mean anything for academic job market? [D]

0 Upvotes
  1. I am applying to academic jobs. We are told to include a section on "impact". I am wondering if the total number of HuggingFace downloads of custom models I have trained would be considered legit impact, or if people would think this was all bots.

  2. Relatedly, for industry (AI labs), is the number of HuggingFace model downloads meaningful?


r/MachineLearning • • 1d ago

Discussion [D] Simple Questions Thread

2 Upvotes

Please post your questions here instead of creating a new thread. Encourage others who create new posts for questions to post here instead!

Thread will stay alive until next one so keep posting after the date in the title.

Thanks to everyone for answering questions in the previous thread!


r/MachineLearning • • 2d ago

Discussion For those who just submit to workshop [D]

69 Upvotes

More recently, I am seeing a lot of posts on the sub regarding the workshop acceptance, with lack of funds to attend the conference. I unfortunately want to just say that workshop papers are not given any importance in the community. I personally consider workshops to either get initial feedback, advertise my work, or just prefer to attend it for the discussions with the community members.

Thought of posting this as I am recently seeing some undegrads/masters students submitting 3-4 papers in the workshops. Even came across a Twitter profile, who had 6-7 workshops in 3 months, and claim to have the PhD worth of work done in three months.

Also please don't expect an explicit funding (apart from organizers, and in some cases your lab may fund it). There's huge problem with the funding, many students even don't it get for main venues. I definitely expect a lot of downvotes on this post, particular coz this is not what many would like to hear, but unfortunately is the reality.


r/MachineLearning • • 1d ago

Discussion WM PAI Workshop at NeurIPS — confused about the acceptance/rejection process [D]

0 Upvotes

I submitted a paper to the WM PAI workshop at NeurIPS and can now see the reviews and decision on OpenReview, but I didn't receive an official acceptance/rejection email.

My reviewer scores were 8, 5, and 4, all with confidence 4, and the paper was ultimately rejected.

What I find a little confusing is that I also reviewed two papers for the same workshop. Both had an average score around 7, but I can see that they were also rejected.

My own submission number was in the 20s, and I submitted on the last day of the submission window, so I assumed there probably weren't a huge number of submissions.

This makes me wonder:

Does anyone know if any papers have actually been accepted to this workshop?

Is it normal for a NeurIPS workshop to reject a large fraction of submissions even with relatively high review scores?

Could the workshop organizers simply be delaying the official notification emails, while the decisions are already visible on OpenReview?

Or am I misunderstanding how the workshop selection process works? 😅

Has anyone else submitted/reviewed for this workshop and received an official decision email?


r/MachineLearning • • 2d ago

Research Tokenization: A Survey for Modern NLP [R]

45 Upvotes

Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field.

We cover every aspect of tokenization: algorithms, evaluations, multilinguality, encodings, theory, etc. We even cover what you might want to replace tokenizers with (e.g., latent or visual tokenization). We also cover some topics that are closely adjacent to tokenization, such as constrained generation, token healing, and tokenizer security concerns.

Check it out!

https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlp


r/MachineLearning • • 2d ago

Research Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models [R]

Thumbnail
gallery
37 Upvotes

I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.

Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.

And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.

The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.


r/MachineLearning • • 2d ago

Research Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes [R]

Thumbnail
gallery
127 Upvotes

coupled-jump.github.io

Hi everyone, I’m happy to share our recent NeurIPS 2026 paper, a collaboration across Google, Google DeepMind and Stony Brook University.

We study a mismatch in joint text and image generation: a model can describe the correct solution to a maze while drawing a different path. Generating both outputs in parallel doesn’t necessarily keep them consistent.

Our sampler, CO₂Jump, uses text confidence and cross-modal attention to guide image updates during sampling. It also allows low-confidence tokens to be masked again and regenerated, so earlier decisions can be revised as generation progresses.

CO₂Jump uses one model forward pass per denoising step. The sampler itself requires no additional training; our experiments compare sampling methods using the same task-specific fine-tuned model.

We evaluate image editing, maze solving and nonograms, and introduce three datasets: JEdit-1M, JMaze-200K and JNono-200K. On the puzzle benchmarks, joint accuracy requires both the textual answer and generated image to be correct. Across 8–512 sampling steps, CO₂Jump was the only sampler we compared that improved monotonically on both editing quality and grounding.

I’d be interested in suggestions for other tasks where text–image consistency and correctness can be evaluated together. Happy to discuss the method, evaluation or limitations.


r/MachineLearning • • 1d ago

Discussion What's up with AAAI round 2 reviews? [D]

0 Upvotes

Has anyone received papers to review for round 2?


r/MachineLearning • • 1d ago

Discussion Does TMLR Confirmation email take time?[D]

0 Upvotes

I can see the submission on Openreview but didn't get any email/notif. It's been 2 hours, I heard I'm supposed to choose action editor or smth. Am I missing something?

First publication ever pls be kind to the noob.


r/MachineLearning • • 2d ago

Project Multi scan radar object classification on RadarScenes [P]

Thumbnail
gallery
4 Upvotes

Hello all,

I built a radar object classifier on RadarScenes, extending a prior single-scan classifier to accumulate observations over a tracked object's history instead of classifying each scan in isolation.

A single RadarScenes object instance contains only about 2.9 radar points on average, very sparse. A single scan also can't capture temporal characteristics: RCS and micro-Doppler both vary continuously as an object moves. Pedestrians produce characteristic micro-Doppler from limb motion; different object classes show different RCS fluctuation patterns as aspect angle and scattering geometry change scan to scan. Accumulating observations gives both higher point density and provides temporal dynamics.

Multi-scan baseline

DeepReflecs encoder (Ulrich, Glaser & Timm, RadarConf 2021), PointNet style, per point shared weights, on single scans across car, large_vehicle, two_wheeler, pedestrian, pedestrian_group: 0.7370 macro F1.

Using RadarScenes' persistent `track_id`, I build a causal, N=20, per track sliding-window buffer:

- x_seq/y_seq: Global, odometry-corrected coordinates recentered per scan on the object centroid. Unlike x_cc/y_cc (car-frame coordinates that accumulate over time to form a trajectory).

- Cross sensor buffer: whichever of the 4 sensors currently observe the track push to the same buffer.

- Stride 1, causal: every new scan updates the buffer and produces a prediction. No future context, real time streaming compatible.

- Each scan is encoded once by a frozen per scan encoder and cached

- Fusion concatenates the causal GRU's hidden state (order aware) with an order-invariant pooled embedding (all N scans' points as one set, no sequence structure) through a small trained mlp head.

Results

Model Macro F1 Delta
Single scan 0.7370 (baseline)
20 scan point pooling 0.8613 +0.1243
Causal GRU 0.8895 +0.0282 over pooling
GRU + pooled embedding (fusion) 0.8897 +0.0002 over GRU, noise

Pooling alone, no sequence model, no notion of scan order at all, recovers +0.1243 macro F1. The GRU adds a real but much smaller +0.0282 on top. Fusion adds nothing measurable beyond the GRU.

Ablation

Llarger GRUs, a Transformer, a state space model, point level self attention, all trained on the exact same frozen per scan embeddings, land inside a 0.86 to 0.89 band, a 0.03 spread. End to end fine tuning of the frozen encoder makes things slightly worse (about -0.002 to -0.003), not better.

Conclusion

In this setup, the largest gain comes from giving the model more observations of the same tracked object: 20-scan point pooling improves macro F1 from 0.7370 to 0.8613 without using scan order at all.

Temporal modelling then provides a further, meaningful improvement. The causal GRU reaches 0.8895, adding +0.0282 over the pooled representation. So temporal ordering clearly contributes useful information; it just accounts for a smaller portion of the overall gain than observation accumulation.

With the per-scan encoder frozen, the different sequence architectures tested, suggests that the quality of the per-scan representation is the bottleneck than the particular mechanism used to aggregate the sequence.

Full report, every ablation, confusion matrix, coordinate frame reasoning: https://github.com/brunopinto900/radar-ml-autonomous-driving/blob/track-accumulation/final_report.md

Thank you.


r/MachineLearning • • 2d ago

Discussion How should I follow up with TMLR submission once all review responses are submitted [D]

5 Upvotes

I have responded to all the reviewers with proper rebuttals and modified draft. One interacted with me and after acknowledging the response kinda disappeared again after asking additional questions. I have replied to his additional questions too but he hasn't raised his score. What will happen if he doesn't reply anymore? Like will AE still consider that an 'Yes' as it was only minor comments which we incorporated in the paper. And one negative reviewer just ghosted us. I have 1/3 explicit positive review currently


r/MachineLearning • • 2d ago

Research NeurIPS SLM Agents or RoboPAD Workshops [D]

0 Upvotes

Hey, just got a papers accepted at both these workshops! Would love to connect with others also attending the workshop


r/MachineLearning • • 2d ago

Project LessThink-Qwen3-4B: the same model, with far less thinking [P]

24 Upvotes

I post-trained Qwen3-4B to spend 44% fewer tokens on reasoning, keeping its knowledge and answer style. The whole pipeline ran on one GPU.

folks, you can check it out on : https://5ivatej.com/lessthink/


r/MachineLearning • • 2d ago

Research Isolation Forest performs best with 1.0 as max_samples [R]

2 Upvotes

I am currently using the dataset CICIDS2017 to train an Isolation Forest model for anomaly detection, I am currently using a split that consists of 70% of ONLY BENIGN traffic for training, 15% BENIGN and 50% attacks for validation and the rest for testing, I used the validation to test with some different max_samples values, # MAX_SAMPLES_VALUES = [256, 4_096, 16_384, 100_000, 200_000, 400_000, 800_000, 1.0] and to calibrate some thresholds that maximixe different statistics (1 for max F1 score, 1 being the closest point to (0,1) on the ROC curve, 5 that set a max limit FPR), the max_samples thing is what is really blowing me off.
At the beginning I only tested up to 200k because seeing that the paper uses 256 as standard value I already thought that 200k was really high, then I saw that a low value is recommended only when anomalies are included in training, which is not my case, so I tested higher values and as it appears 1.0 (max value) is effectively the best performing with ~94% recall and ~7.6% FPR on the ROC threshold, while at 200k I get ~91% recall and ~10% FPR, all this with only 30 secs more in training.

I also executed a cross-dataset test with CSE-CIC-IDS2018, the performance is basically the same (bad) both with 200k and 1.0, my question then is, should I keep 1.0 as operational max_samples value or should I decrease it back to 200k/400k?

I forgot to mention, is my training approach bad? I've read from multiple sources that someone trains on whole dataset, someone adds anomalies to training set, etc..., is what i'm doing okay or should I change approach? I think so because this is one of the cases where clusters of anomalies can cause swamping/masking, and we have high-dimensional data, etc..