sisterdong

sisterdong

4 followers · 6 following

Block or Report

Block or report sisterdong

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Lists (1)

Sort

👀 Interviews

2 repositories

Beta Lists are currently in beta. Share feedback and report bugs.

Stars

rom1504 / cc2dataset

Easily convert common crawl to a dataset of caption and document. Image/text Audio/text Video/text, ...

Python 301 23 Updated Dec 9, 2023

DmitryRyumin / CVPR-2023-24-Papers

CVPR 2023-2024 Papers: Dive into advanced research presented at the leading computer vision conference. Keep up to date with the latest developments in computer vision and deep learning. Code inclu…

Python 333 22 Updated Jun 28, 2024

pymupdf / PyMuPDF-Utilities

Demos, examples and utilities using PyMuPDF

Jupyter Notebook 504 140 Updated Jul 1, 2024

EleutherAI / gpt-neox

An implementation of model parallel autoregressive transformers on GPUs, based on the Megatron and DeepSpeed libraries

Python 6,694 977 Updated Jun 28, 2024

EleutherAI / pythia

The hub for EleutherAI's work on interpretability and learning dynamics

Jupyter Notebook 2,123 155 Updated Jun 18, 2024

NielsRogge / Vision-Transformer-papers

This repository contains an overview of important follow-up works based on the original Vision Transformer (ViT) by Google.

127 9 Updated Jan 3, 2022

VikParuchuri / marker

Convert PDF to markdown quickly with high accuracy

Python 13,769 683 Updated Jun 30, 2024

attardi / wikiextractor

A tool for extracting plain text from Wikipedia dumps

Python 3,681 958 Updated May 23, 2024

graphistry / pygraphistry

PyGraphistry is a Python library to quickly load, shape, embed, and explore big graphs with the GPU-accelerated Graphistry visual graph analyzer

Python 2,090 206 Updated Jul 1, 2024

Visualize-ML / Book6_First-Course-in-Data-Science

Book_6_《数据有道》 | 鸢尾花书：从加减乘除到机器学习；欢迎大家批评指正！纠错多的同学会得到赠书感谢！

Jupyter Notebook 1,678 316 Updated Jun 29, 2024

jawah / charset_normalizer

Truly universal encoding detector in pure Python

Python 537 49 Updated Jul 1, 2024

huggingface / datatrove

Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.

Python 1,735 103 Updated Jun 28, 2024

OpenMatch / NeuScraper

[ACL 2024] This is the code repo for our ACL’24 paper "Cleaner Pretraining Corpus Curation with Neural Web Scraping".

Python 196 15 Updated Jun 20, 2024

princeton-nlp / SWE-bench

[ICLR 2024] SWE-Bench: Can Language Models Resolve Real-world Github Issues?

Python 1,440 238 Updated Jul 1, 2024

openai / transformer-debugger

Python 3,965 231 Updated Jun 4, 2024

InflectionAI / Inflection-Benchmarks

Public Inflection Benchmarks

67 2 Updated Mar 6, 2024

LDNOOBW / List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words

List of Dirty, Naughty, Obscene, and Otherwise Bad Words

2,819 654 Updated Jun 19, 2024

gkamradt / LLMTest_NeedleInAHaystack

Doing simple retrieval from LLM models at various context lengths to measure accuracy

Jupyter Notebook 1,276 131 Updated Jun 20, 2024

HqWu-HITCS / Awesome-LLM-Survey

An Awesome Collection for LLM Survey

249 24 Updated May 2, 2024

ksOAn6g5 / TaiSu

TaiSu（太素）--a large-scale Chinese multimodal dataset（亿级大规模中文视觉语言预训练数据集）

Python 164 10 Updated Nov 17, 2023

NiuTrans / Classical-Modern

非常全的文言文（古文）-现代文平行语料

Python 952 205 Updated Apr 21, 2024

kpu / kenlm

KenLM: Faster and Smaller Language Model Queries

C++ 2,438 511 Updated Feb 25, 2024

JessicaTegner / pypandoc

Thin wrapper for "pandoc" (MIT)

Python 834 108 Updated Jun 4, 2024

Zhen-Tan-dmml / LLM4Annotation

188 9 Updated Mar 9, 2024

oscar-project / ungoliant

🕷️ The pipeline for the OSCAR corpus

Rust 154 14 Updated Dec 18, 2023

chatnoir-eu / web-content-extraction-benchmark

Web Content Extraction Benchmark

Python 14 1 Updated May 24, 2024

adbar / trafilatura

Python & command-line tool to gather text on the Web: Crawling & scraping, content extraction, metadata. TXT, Markdown, CSV & XML output.

Python 3,171 239 Updated Jul 2, 2024

modin-project / modin

Modin: Scale your Pandas workflows by changing a single line of code

Python 9,586 648 Updated Jun 28, 2024

Ethan-yt / guwenbert

GuwenBERT: 古文预训练语言模型（古文BERT） A Pre-trained Language Model for Classical Chinese (Literary Chinese)

480 41 Updated Aug 31, 2021

chujiezheng / chat_templates

Chat Templates for 🤗 HuggingFace Large Language Models

Jinja 322 31 Updated Jun 27, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly