Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
27
Follow
Fred23456789's profile picture
armant1372's profile picture
Fishtiks's profile picture
94 followers
·
128 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
replied
to
their
post
about 2 hours ago
Say hello to the AlephLLM: Mini-Beatrix - in her huggingface space! She is currently stepped at 8000 steps in the first couple datasets, so she's not very smart yet. https://huggingface.co/spaces/AbstractPhil/alephllm-chat Be warned, whatever you say WILL be recorded in a public cache, WHEN the chat version works. For now she records nothing. The idea is to help debug the K/V cache, and I would rather the data accumulated be shared. If you wish to speak to her in private I will include a toggle, that way you'll see that nothing is recorded when you speak and you can still have a private chat with her. For now she's simply auto-completing, so have fun with her. The AlephLLM prototype is currently in full training with SDPA attention. https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype. As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults. It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning. Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded.
posted
an
update
about 7 hours ago
Say hello to the AlephLLM: Mini-Beatrix - in her huggingface space! She is currently stepped at 8000 steps in the first couple datasets, so she's not very smart yet. https://huggingface.co/spaces/AbstractPhil/alephllm-chat Be warned, whatever you say WILL be recorded in a public cache, WHEN the chat version works. For now she records nothing. The idea is to help debug the K/V cache, and I would rather the data accumulated be shared. If you wish to speak to her in private I will include a toggle, that way you'll see that nothing is recorded when you speak and you can still have a private chat with her. For now she's simply auto-completing, so have fun with her. The AlephLLM prototype is currently in full training with SDPA attention. https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype. As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults. It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning. Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded.
updated
a model
about 7 hours ago
AbstractPhil/alephllm-mini-beatrix-training
View all activity
Organizations
AbstractPhil
's models
213
Sort: Recently updated
AbstractPhil/geolip-scene-classifier-proto
Updated
Mar 1
AbstractPhil/geometric-experiment-history
Updated
Feb 25
AbstractPhil/sd15-geovocab-lora-prototype
Text-to-Image
•
Updated
Feb 21
AbstractPhil/geovae-proto
Updated
Feb 21
•
1
AbstractPhil/geovocab-patch-maker
Updated
Feb 21
•
6
AbstractPhil/grid-geometric-multishape
Updated
Feb 20
AbstractPhil/grid-geometric-classifier-sliding-proto
Image Classification
•
Updated
Feb 20
AbstractPhil/grid-geometric-classifier-proto
Other
•
Updated
Feb 13
AbstractPhil/sd15-geoflow-test-44-1000
Text-to-Image
•
Updated
Feb 8
AbstractPhil/sd15-geoflow-test-44-10_000
Text-to-Image
•
Updated
Feb 8
AbstractPhil/sd15-geoflow-characters
Text-to-Image
•
Updated
Feb 7
AbstractPhil/sd15-rectified-geometric-matching
Text-to-Image
•
Updated
Feb 7
•
1
AbstractPhil/sd15-geoflow-object-association
Text-to-Image
•
Updated
Feb 7
•
1
AbstractPhil/ksimplex-llm-prototype
Updated
Feb 4
•
4
•
1
AbstractPhil/ksimplex-linear-experiments
Updated
Feb 4
AbstractPhil/mobiusnet-distillations
Image Classification
•
Updated
Feb 3
AbstractPhil/sbert-voynich-translation
Feature Extraction
•
Updated
Feb 3
AbstractPhil/tiny-flux-deep
Text-to-Image
•
0.2B
•
Updated
Feb 1
•
154
•
5
AbstractPhil/tinyflux-lailah-loras
Updated
Jan 31
AbstractPhil/tinyflux-test-e2e
Updated
Jan 29
•
2
AbstractPhil/tinyflux-experts
Updated
Jan 28
•
2
AbstractPhil/sd15-flow-matching
Text-to-Image
•
Updated
Jan 27
•
219
•
3
AbstractPhil/tiny-flux
Text-to-Image
•
10.7M
•
Updated
Jan 22
•
60
•
3
AbstractPhil/mobiusnet-collective
Updated
Jan 11
AbstractPhil/mobiusnet
Updated
Jan 10
AbstractPhil/vit-beatrix-contrarian
Updated
Dec 29, 2025
AbstractPhil/vit-beatrix-contrarian-baselines
Updated
Dec 29, 2025
AbstractPhil/beatrix-diffusion-proto
Updated
Dec 28, 2025
•
16
AbstractPhil/global_fractal_router
Updated
Dec 13, 2025
AbstractPhil/agatha-diffusion-proto
Updated
Dec 9, 2025
Previous
1
2
3
4
5
6
...
8
Next