DFCI-Hale

non-profit
Verified

AI & ML interests

This organization serves as a platform for benchmarking classic and novel 3D segmentation models to facilitate our research in DFCI-AIOS on Next-generation organ segmentation model development.

Parveshiiiiย 
posted an update about 1 month ago
view post
Post
540
๐Ÿš€ Sonic: A lightweight Python audio processing library with tempo matching, BPM detection, time-stretching, resampling & track blending โ€” now with GPU (CUDA) acceleration for 10x speed!

Perfect for quick remixes, batch edits or syncing tracks.

๐Ÿ‘‰ https://github.com/Parveshiiii/Sonic

#Python #AudioProcessing #OpenSource #PyTorch
Parveshiiiiย 
posted an update about 1 month ago
view post
Post
1619
Excited to announce my latest open-source release on Hugging Face: Parveshiiii/breast-cancer-detector.

This model has been trained and validated on external datasets to support medical research workflows. It is designed to provide reproducible benchmarks and serve as a foundation for further exploration in healthcare AI.

Key highlights:
- Built for medical research and diagnostic study contexts
- Validated against external datasets for reliability
- Openly available to empower the community in building stronger, more effective solutions

This release is part of my ongoing effort to make impactful AI research accessible through **Modotte**. A detailed blog post explaining the methodology, dataset handling, and validation process will be published soon.

You can explore the model here: Parveshiiii/breast-cancer-detector

#AI #MedicalResearch #DeepLearning #Healthcare #OpenSource #HuggingFace

Parveshiiiiย 
posted an update about 2 months ago
view post
Post
2933
Just did something Iโ€™ve been meaning to try for ages.

In only 3 hours, on 10 billion+ tokens, I trained a custom BPE + tiktoken-style tokenizer using my new library microtok โ€” and it hits the same token efficiency as Qwen3.

Tokenizers have always felt like black magic to me. We drop them into every LLM project, but actually training one from scratch? That always seemed way too complicated.

Turns out it doesnโ€™t have to be.

microtok makes the whole process stupidly simple โ€” literally just 3 lines of code. No heavy setup, no GPU required. I built it on top of the Hugging Face tokenizers library so it stays clean, fast, and actually understandable.

If youโ€™ve ever wanted to look under the hood and build your own optimized vocabulary instead of just copying someone elseโ€™s, this is the entry point youโ€™ve been waiting for.

I wrote up the full story, threw in a ready-to-run Colab template, and dropped the trained tokenizer on Hugging Face.

Blog โ†’ https://parveshiiii.github.io/blogs/microtok/
Trained tokenizer โ†’ https://huggingface.co/Parveshiiii/microtok
GitHub repo โ†’ https://github.com/Parveshiiii/microtok
Parveshiiiiย 
posted an update 3 months ago
view post
Post
343
Introducing Seekify โ€” a truly nonโ€‘rateโ€‘limiting search library for Python

Tired of hitting rate limits when building search features? Iโ€™ve built Seekify, a lightweight Python library that lets you perform searches without the usual throttling headaches.

๐Ÿ”น Key highlights

- Simple API โ€” plug it in and start searching instantly

- No rateโ€‘limiting restrictions

- Designed for developers who need reliable search in projects, scripts, or apps

๐Ÿ“ฆ Available now on PyPI:

pip install seekify

๐Ÿ‘‰ Check out the repo: https:/github.com/Parveshiiii/Seekify
Iโ€™d love feedback, contributions, and ideas for realโ€‘world use cases. Letโ€™s make search smoother together!
Parveshiiiiย 
posted an update 4 months ago
view post
Post
1644
๐Ÿš€ Wanna train your own AI Model or Tokenizer from scratch?

Building models isnโ€™t just for big labs anymore โ€” with the right data, compute, and workflow, you can create **custom AI models** and **tokenizers** tailored to any domain. Whether itโ€™s NLP, domainโ€‘specific datasets, or experimental architectures, training from scratch gives you full control over vocabulary, embeddings, and performance.

โœจ Why train your own?
- Full control over vocabulary & tokenization
- Domainโ€‘specific optimization (medical, legal, technical, etc.)
- Better performance on niche datasets
- Freedom to experiment with architectures

โšก The best part?
- Tokenizer training (TikToken / BPE) can be done in **just 3 lines of code**.
- Model training runs smoothly on **Google Colab notebooks** โ€” no expensive hardware required.

๐Ÿ“‚ Try out my work:
- ๐Ÿ”— https://github.com/OE-Void/Tokenizer-from_scratch
- ๐Ÿ”— https://github.com/OE-Void/GPT
Parveshiiiiย 
posted an update 4 months ago
view post
Post
266
๐Ÿ“ข The Announcement
Subject: XenArcAI is now Modotte โ€“ A New Chapter Begins! ๐Ÿš€

Hello everyone,

We are thrilled to announce that XenArcAI is officially rebranding to Modotte!

Since our journey began, weโ€™ve been committed to pushing the boundaries of AI through open-source innovation, research, and high-quality datasets. As we continue to evolve, we wanted a name that better represents our vision for a modern, interconnected future in the tech space.

What is changing?

The Name: Moving forward, all our projects, models, and community interactions will happen under the Modotte banner.

The Look: Youโ€™ll see our new logo and a fresh color palette appearing across our platforms.

What is staying the same?

The Core Team: Itโ€™s still the same people behind the scenes, including our founder, Parvesh Rawal.

Our Mission: We remain dedicated to releasing state-of-the-art open-source models and datasets.

Our Continuity: All existing models, datasets, and projects will remain exactly as they areโ€”just with a new home.

This isnโ€™t just a change in appearance; itโ€™s a commitment to our next chapter of growth and discovery. We are so grateful for your ongoing support as we step into this new era.

Welcome to the future. Welcome to Modotte.

Best regards, The Modotte Team
Parveshiiiiย 
posted an update 5 months ago
view post
Post
3599
Hey everyone!
Weโ€™re excited to introduce our new Telegram group: https://t.me/XenArcAI

This space is built for **model builders, tech enthusiasts, and developers** who want to learn, share, and grow together. Whether youโ€™re just starting out or already deep into AI/ML, youโ€™ll find a supportive community ready to help with knowledge, ideas, and collaboration.

๐Ÿ’ก Join us to:
- Connect with fellow developers and AI enthusiasts
- Share your projects, insights, and questions
- Learn from others and contribute to a growing knowledge base

๐Ÿ‘‰ If youโ€™re interested, hop in and be part of the conversation: https://t.me/XenArcAI
  • 12 replies
ยท
Parveshiiiiย 
posted an update 6 months ago
view post
Post
1672
Another banger from XenArcAI! ๐Ÿ”ฅ

Weโ€™re thrilled to unveil three powerful new releases that push the boundaries of AI research and development:

๐Ÿ”— https://huggingface.co/XenArcAI/SparkEmbedding-300m

- A lightning-fast embedding model built for scale.
- Optimized for semantic search, clustering, and representation learning.

๐Ÿ”— https://huggingface.co/datasets/XenArcAI/CodeX-7M-Non-Thinking

- A massive dataset of 7 million code samples.
- Designed for training models on raw coding patterns without reasoning layers.

๐Ÿ”— https://huggingface.co/datasets/XenArcAI/CodeX-2M-Thinking

- A curated dataset of 2 million code samples.
- Focused on reasoning-driven coding tasks, enabling smarter AI coding assistants.

Together, these projects represent a leap forward in building smarter, faster, and more capable AI systems.

๐Ÿ’ก Innovation meets dedication.
๐ŸŒ Knowledge meets responsibility.


Parveshiiiiย 
posted an update 6 months ago
view post
Post
3066
SparkEmbedding - SoTA cross lingual retrieval

Iam very happy to announce our latest embedding model sparkembedding-300m base on embeddinggemma-300m we fine tuned it on 1m extra examples spanning over 119 languages and result is this model achieves exceptional cross lingual retrieval

Model: https://huggingface.co/XenArcAI/SparkEmbedding-300m
Parveshiiiiย 
posted an update 7 months ago
view post
Post
232
AIRealNet - SoTA - Image detection model

Weโ€™re proud to release AIRealNet โ€” a binary image classifier built to detect whether an image is AI-generated or a real human photograph. Based on SwinV2 and fine-tuned on the AI-vs-Real dataset, this model is optimized for high-accuracy classification across diverse visual domains.

If you care about synthetic media detection or want to explore the frontier of AI vs human realism, weโ€™d love your support. Please like the model and try it out. Every download helps us improve and expand future versions.

Model page: https://huggingface.co/XenArcAI/AIRealNet
Parveshiiiiย 
posted an update 7 months ago
view post
Post
4512
Ever wanted an openโ€‘source deep research agent? Meet Deepresearchโ€‘Agent ๐Ÿ”๐Ÿค–

1. Multiโ€‘step reasoning: Reflects between steps, fills gaps, iterates until evidence is solid.

2. Researchโ€‘augmented: Generates queries, searches, synthesizes, and cites sources.

3. Fullstack + LLMโ€‘friendly: React/Tailwind frontend, LangGraph/FastAPI backend; works with OpenAI/Gemini.


๐Ÿ”— GitHub: https://github.com/Parveshiiii/Deepresearch-Agent
Parveshiiiiย 
posted an update 8 months ago
view post
Post
3127
๐Ÿš€ Big news from XenArcAI!

Weโ€™ve just released our new dataset: **Bhagwatโ€‘Gitaโ€‘Infinity** ๐ŸŒธ๐Ÿ“–

โœจ Whatโ€™s inside:
- Verseโ€‘aligned Sanskrit, Hindi, and English
- Clean, structured, and ready for ML/AI projects
- Perfect for research, education, and openโ€‘source exploration

๐Ÿ”— Hugging Face: https://huggingface.co/datasets/XenArcAI/Bhagwat-Gita-Infinity

Letโ€™s bring timeless wisdom into modern AI together ๐Ÿ™Œ
Parveshiiiiย 
posted an update 8 months ago
view post
Post
2473
๐Ÿš€ New Release from XenArcAI
Weโ€™re excited to introduce AIRealNet โ€” our SwinV2โ€‘based image classifier built to distinguish between artificial and real images.

โœจ Highlights:
- Backbone: SwinV2
- Input size: 256ร—256
- Labels: artificial vs. real
- Performance: Accuracy 0.999 | F1 0.999 | Val Loss 0.0063

This model is now live on Hugging Face:
๐Ÿ‘‰ https://huggingface.co/XenArcAI/AIRealNet

We built AIRealNet to push forward openโ€‘source tools for authenticity detection, and we canโ€™t wait to see how the community uses it.
Parveshiiiiย 
posted an update 9 months ago
view post
Post
1123
๐Ÿš€ Just Dropped: MathX-5M โ€” Your Gateway to Math-Savvy GPTs

๐Ÿ‘จโ€๐Ÿ”ฌ Wanna fine-tune your own GPT for math?
๐Ÿง  Building a reasoning agent that actually *thinks*?
๐Ÿ“Š Benchmarking multi-step logic across domains?

Say hello to [**MathX-5M**](https://huggingface.co/datasets/XenArcAI/MathX-5M) โ€” a **5 million+ sample** dataset crafted for training and evaluating math reasoning models at scale.

Built by **XenArcAI**, itโ€™s optimized for:
- ๐Ÿ” Step-by-step reasoning with , , and formats
- ๐Ÿงฎ Coverage from arithmetic to advanced algebra and geometry
- ๐Ÿงฐ Plug-and-play with Gemma, Qwen, Mistral, and other open LLMs
- ๐Ÿงต Compatible with Harmony, Alpaca, and OpenChat-style instruction formats

Whether you're prototyping a math tutor, testing agentic workflows, or just want your GPT to solve equations like a proโ€”**MathX-5M is your launchpad**.

๐Ÿ”— Dive in: (https://huggingface.co/datasets/XenArcAI/MathX-5M)

Letโ€™s make open-source models *actually* smart at math.
#FineTuneYourGPT #MathX5M #OpenSourceAI #LLM #XenArcAI #Reasoning #Gemma #Qwen #Mistral

Parveshiiiiย 
posted an update 10 months ago
view post
Post
1106
๐Ÿš€ Launch Alert: Dev-Stack-Agents
Meet your 50-agent senior AI team โ€” principal-level experts in engineering, AI, DevOps, security, product, and more โ€” all bundled into one modular repo.

+ Code. Optimize. Scale. Secure.
- Full-stack execution, Claude-powered. No human bottlenecks.


๐Ÿ”ง Built for Claude Code
Seamlessly plug into Claudeโ€™s dev environment:

* ๐Ÿง  Each .md file = a fully defined expert persona
* โš™๏ธ Claude indexes them as agents with roles, skills & strategy
* ๐Ÿค– You chat โ†’ Claude auto-routes to the right agent(s)
* โœ๏ธ Want precision? Just call @agent-name directly
* ๐Ÿ‘ฅ Complex task? Mention multiple agents for team execution

Examples:

"@security-auditor please review auth flow for risks"
"@cloud-architect + @devops-troubleshooter โ†’ design a resilient multi-region setup"
"@ai-engineer + @legal-advisor โ†’ build a privacy-safe RAG pipeline"


๐Ÿ”— https://github.com/Parveshiiii/Dev-Stack-Agents
MIT License | Claude-Ready | PRs Welcome

Parveshiiiiย 
posted an update 10 months ago
view post
Post
2715
๐Ÿง  Glimpses of AGI โ€” A Vision for All Humanity
What if AGI wasnโ€™t just a distant dreamโ€”but a blueprint already unfolding?

Iโ€™ve just published a deep dive called Glimpses of AGI, exploring how scalable intelligence, synthetic reasoning, and alignment strategies are paving a new path forward. This isnโ€™t your average tech commentaryโ€”itโ€™s a bold vision for conscious AI systems that reason, align, and adapt beyond narrow tasks.

๐Ÿ” Read it, upvote it if it sparks something, and letโ€™s ignite a collective conversation about the future of AGI.

https://huggingface.co/blog/Parveshiiii/glimpses-of-agi


Parveshiiiiย 
posted an update 11 months ago
view post
Post
2868
๐Ÿง  MathX-5M by XenArcAI โ€” Scalable Math Reasoning for Smarter LLMs

Introducing MathX-5M, a high-quality, instruction-tuned dataset built to supercharge mathematical reasoning in large language models. With 5 million rigorously filtered examples, it spans everything from basic arithmetic to advanced calculusโ€”curated from public sources and enhanced with synthetic data.

๐Ÿ” Key Highlights:
- Step-by-step reasoning with verified answers
- Covers algebra, geometry, calculus, logic, and more
- RL-validated correctness and multi-stage filtering
- Ideal for fine-tuning, benchmarking, and educational AI

๐Ÿ“‚ - https://huggingface.co/datasets/XenArcAI/MathX-5M


  • 1 reply
ยท
Threatthriverย 
posted an update 12 months ago
view post
Post
319
New Dataset Released