Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
27
Follow
CosmicCrafter's profile picture
Fred23456789's profile picture
juiceb0xc0de's profile picture
94 followers
·
128 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
replied
to
their
post
about 5 hours ago
Say hello to the AlephLLM: Mini-Beatrix - in her huggingface space! She is currently stepped at 8000 steps in the first couple datasets, so she's not very smart yet. https://huggingface.co/spaces/AbstractPhil/alephllm-chat Be warned, whatever you say WILL be recorded in a public cache, WHEN the chat version works. For now she records nothing. The idea is to help debug the K/V cache, and I would rather the data accumulated be shared. If you wish to speak to her in private I will include a toggle, that way you'll see that nothing is recorded when you speak and you can still have a private chat with her. For now she's simply auto-completing, so have fun with her. The AlephLLM prototype is currently in full training with SDPA attention. https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype. As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults. It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning. Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded.
posted
an
update
about 10 hours ago
Say hello to the AlephLLM: Mini-Beatrix - in her huggingface space! She is currently stepped at 8000 steps in the first couple datasets, so she's not very smart yet. https://huggingface.co/spaces/AbstractPhil/alephllm-chat Be warned, whatever you say WILL be recorded in a public cache, WHEN the chat version works. For now she records nothing. The idea is to help debug the K/V cache, and I would rather the data accumulated be shared. If you wish to speak to her in private I will include a toggle, that way you'll see that nothing is recorded when you speak and you can still have a private chat with her. For now she's simply auto-completing, so have fun with her. The AlephLLM prototype is currently in full training with SDPA attention. https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype. As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults. It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning. Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded.
updated
a model
about 10 hours ago
AbstractPhil/alephllm-mini-beatrix-training
View all activity
Organizations
AbstractPhil
's datasets
83
Sort: Recently updated
AbstractPhil/alephllm-chat-history
Viewer
•
Updated
about 12 hours ago
•
1
•
14
AbstractPhil/captionbert-8192-v2-consensus
Updated
11 days ago
•
227
AbstractPhil/conceptual-captions-12m-webdataset-berts
Viewer
•
Updated
12 days ago
•
32.3M
•
504
•
1
AbstractPhil/bulk-cc12m-features
Viewer
•
Updated
13 days ago
•
121M
•
3.87k
AbstractPhil/tower-probes-results
Viewer
•
Updated
22 days ago
•
17
•
145
AbstractPhil/qwen-deepfashion-fused
Viewer
•
Updated
Jul 12
•
122k
•
1.2k
•
1
AbstractPhil/qwen-synth-characters-fused
Viewer
•
Updated
Jul 10
•
42.7k
•
933
AbstractPhil/qwen-synth-characters-100-json-test
Viewer
•
Updated
Jul 10
•
1k
•
53
AbstractPhil/anima-brent-90k-cache
Updated
Jul 5
•
37
AbstractPhil/qwen-synth-characters
Viewer
•
Updated
Jul 3
•
61k
•
126
AbstractPhil/qwen-deepfashion
Viewer
•
Updated
Jul 3
•
160k
•
189
AbstractPhil/diffusion-pipe-cache-test1
Viewer
•
Updated
Jun 27
•
8.92k
•
64
AbstractPhil/anima-90k-cache
Updated
Jun 26
•
62
AbstractPhil/diffusion-pretrain-set-ft1
Viewer
•
Updated
Jun 23
•
1.46M
•
1.8k
•
1
AbstractPhil/diffusion-pretrain-set-ft1-1024
Viewer
•
Updated
Jun 11
•
1.14M
•
693
AbstractPhil/sdxl-qwen-phase1-cache
Viewer
•
Updated
Jun 6
•
86k
•
330
AbstractPhil/geolip-sdxl-fid-scoring
Viewer
•
Updated
Jun 5
•
2.8k
•
94
AbstractPhil/sdxl-qwen-phase0
Viewer
•
Updated
Jun 4
•
86k
•
256
•
3
AbstractPhil/IMDB-PUBLIC-SCRAPED
Preview
•
Updated
May 19
•
112
•
1
AbstractPhil/ldhnam-deepfashion_controlnet
Viewer
•
Updated
May 19
•
26k
•
25
AbstractPhil/ffhq_flux_latents_repaired
Viewer
•
Updated
May 19
•
40.8k
•
213
AbstractPhil/synthetic-characters
Viewer
•
Updated
May 19
•
149k
•
345
AbstractPhil/CN_pose3D_V10_512
Viewer
•
Updated
May 19
•
66.5k
•
63
AbstractPhil/CN_pose3D_V7_512
Viewer
•
Updated
May 19
•
255k
•
453
AbstractPhil/synthetic-object-relations-json
Viewer
•
Updated
May 18
•
5k
•
39
AbstractPhil/cc-task1-json
Preview
•
Updated
May 18
•
54
AbstractPhil/cc-prompts-sharded
Viewer
•
Updated
May 15
•
3.32M
•
10
AbstractPhil/json-coco-format
Viewer
•
Updated
May 14
•
129k
•
174
AbstractPhil/svae-freckles-4096-cifar10
Viewer
•
Updated
Apr 10
•
60k
•
64
AbstractPhil/ryan-spearman-prepared-features
Viewer
•
Updated
Mar 27
•
1
•
70
Previous
1
2
3
Next