Skip to content

Refactor RegexTokenizer to handle empty stats and update max function usage - #17

Open
Conceptaniel wants to merge 4 commits into
ImadSaddik:mainfrom
Conceptaniel:main
Open

Refactor RegexTokenizer to handle empty stats and update max function usage#17
Conceptaniel wants to merge 4 commits into
ImadSaddik:mainfrom
Conceptaniel:main

Conversation

@Conceptaniel

Copy link
Copy Markdown
  • Added a check to break the loop if stats is empty in RegexTokenizer.
  • Updated the way the pair with the highest count is determined using max with items.

Update execution counts in DataCleaning and BytePairEncoding notebooks

  • Adjusted execution counts to maintain consistency after code changes.
  • Updated Python version in notebook metadata.

Fix execution counts and add new cell in TransformerModel notebook

  • Corrected execution counts for several cells to reflect the current order.
  • Added a new empty code cell for future use.
  • Updated Python version in notebook metadata.

Add read_plain_chat function to DataCleaning notebook

  • Introduced a new function to clean and parse plain text chat exports.
  • The function handles various message formats and normalizes the output into a DataFrame.

Conceptaniel and others added 3 commits July 21, 2026 11:32
… usage

- Added a check to break the loop if stats is empty in RegexTokenizer.
- Updated the way the pair with the highest count is determined using max with items.

Update execution counts in DataCleaning and BytePairEncoding notebooks

- Adjusted execution counts to maintain consistency after code changes.
- Updated Python version in notebook metadata.

Fix execution counts and add new cell in TransformerModel notebook

- Corrected execution counts for several cells to reflect the current order.
- Added a new empty code cell for future use.
- Updated Python version in notebook metadata.

Add read_plain_chat function to DataCleaning notebook

- Introduced a new function to clean and parse plain text chat exports.
- The function handles various message formats and normalizes the output into a DataFrame.
@ImadSaddik

Copy link
Copy Markdown
Owner

Hi @Conceptaniel,

Thank you for contributing to this project.

I will take a look at your PR this weekend.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants