cpython

Commit Graph

Author	SHA1	Message	Date
Lysandros Nikolaou	d66bc9e8a7	gh-107967: Fix infinite recursion on invalid escape sequence warning (#107968 )	2023-08-15 11:26:42 +00:00
Pablo Galindo Salgado	da8f87b7ea	gh-107015: Remove async_hacks from the tokenizer (#107018 )	2023-07-26 16:34:15 +01:00
Menelaos Kotoglou	76e20c361c	gh-106989: Remove tok report warnings (#106993 )	2023-07-22 14:23:23 +02:00
Inada Naoki	d5bd32fb48	gh-104922: remove PY_SSIZE_T_CLEAN (#106315 )	2023-07-02 15:07:46 +09:00
Lysandros Nikolaou	6586cee27f	gh-105938: Emit a SyntaxWarning for escaped braces in an f-string (#105939 )	2023-06-20 12:38:46 +00:00
Lysandros Nikolaou	d382ad4915	gh-105820: Fix tok_mode expression buffer in file & readline tokenizer (#105828 )	2023-06-15 16:21:24 +00:00
Lysandros Nikolaou	abfbab6415	gh-105718: Fix buffer allocation in tokenizer with readline (#105728 )	2023-06-13 16:18:11 +01:00
Pablo Galindo Salgado	b047fa5e56	gh-105549: Tokenize separately NUMBER and NAME tokens and allow 0-prefixed literals (#105555 )	2023-06-09 21:39:01 +01:00
Pablo Galindo Salgado	c0a6ed3934	gh-105259: Ensure we don't show newline characters for trailing NEWLINE tokens (#105364 )	2023-06-06 12:52:16 +01:00
Lysandros Nikolaou	70f315c2d6	gh-105042: Disable unmatched parens syntax error in python tokenize (#105061 )	2023-05-30 22:52:52 +01:00
Pablo Galindo Salgado	9216e69a87	gh-105069: Add a readline-like callable to the tokenizer to consume input iteratively (#105070 )	2023-05-30 22:43:34 +01:00
Marta Gómez Macías	96fff35325	gh-105017: Include CRLF lines in strings and column numbers (#105030 ) Co-authored-by: Pablo Galindo <pablogsal@gmail.com>	2023-05-28 15:15:53 +01:00
Marta Gómez Macías	86d8f48935	gh-105017: Fix including additional NL token when using CRLF (#105022 ) Co-authored-by: Pablo Galindo Salgado <Pablogsal@gmail.com>	2023-05-27 16:50:43 +00:00
Petr Vaněk	6e62eb2e70	Fix indentation in Parser/tokenizer.c (#105012 )	2023-05-27 12:41:50 +01:00
Lysandros Nikolaou	c90a862cdc	gh-104866: Tokenize should emit NEWLINE after exiting block with comment (#104870 )	2023-05-24 17:18:17 +01:00
Cristián Maureira-Fredes	0a7796052a	gh-102856: Allow comments inside multi-line f-string expresions (#104006 )	2023-05-22 10:30:07 +00:00
Serhiy Storchaka	f3466bc040	gh-98836: Extend PyUnicode_FromFormat() (GH-98838) * Support for conversion specifiers o (octal) and X (uppercase hexadecimal). * Support for length modifiers j (intmax_t) and t (ptrdiff_t). * Length modifiers are now applied to all integer conversions. * Support for wchar_t C strings (%ls and %lV). * Support for variable width and precision (). Support for flag - (left alignment).	2023-05-22 00:32:39 +03:00
Marta Gómez Macías	6715f91edc	gh-102856: Python tokenizer implementation for PEP 701 (#104323 ) This commit replaces the Python implementation of the tokenize module with an implementation that reuses the real C tokenizer via a private extension module. The tokenize module now implements a compatibility layer that transforms tokens from the C tokenizer into Python tokenize tokens for backward compatibility. As the C tokenizer does not emit some tokens that the Python tokenizer provides (such as comments and non-semantic newlines), a new special mode has been added to the C tokenizer mode that currently is only used via the extension module that exposes it to the Python layer. This new mode forces the C tokenizer to emit these new extra tokens and add the appropriate metadata that is needed to match the old Python implementation. Co-authored-by: Pablo Galindo <pablogsal@gmail.com>	2023-05-21 01:03:02 +01:00
Pablo Galindo Salgado	ff7f731632	gh-104658: Fix location of unclosed quote error for multiline f-strings (#104660 )	2023-05-20 14:07:05 +01:00
Hugo van Kemenade	d513ddee94	Trim trailing whitespace and test on CI (#104275 ) Co-authored-by: Alex Waygood <Alex.Waygood@Gmail.com>	2023-05-08 17:03:52 +03:00
Pablo Galindo Salgado	eba64d2afb	gh-104169: Ensure the tokenizer doesn't overwrite previous errors (#104170 )	2023-05-04 15:15:26 +01:00
Lysandros Nikolaou	ef0df5284f	gh-97556: Raise null bytes syntax error upon null in multiline string (GH-104136)	2023-05-04 14:26:23 +02:00
jx124	5078eedc5b	gh-104016: Fixed off by 1 error in f string tokenizer (#104047 ) Co-authored-by: sunmy2019 <59365878+sunmy2019@users.noreply.github.com> Co-authored-by: Ken Jin <kenjin@python.org> Co-authored-by: Pablo Galindo <pablogsal@gmail.com>	2023-05-01 19:15:47 +00:00
chgnrdv	d5a97074d2	gh-103824: fix use-after-free error in Parser/tokenizer.c (#103993 )	2023-05-01 15:26:43 +00:00
Lysandros Nikolaou	9169a56fad	gh-103656: Transfer f-string buffers to parser to avoid use-after-free (GH-103896) Co-authored-by: Pablo Galindo <pablogsal@gmail.com>	2023-04-27 01:33:31 +00:00
Lysandros Nikolaou	57f8f9a66d	gh-103718: Correctly set f-string buffers in all cases (GH-103815) Turns out we always need to remember/restore fstring buffers in all of the stack of tokenizer modes, cause they might change to `TOK_REGULAR_MODE` and have newlines inside the braces (which is when we need to reallocate the buffer and restore the fstring ones).	2023-04-25 01:31:21 +00:00
Lysandros Nikolaou	cb157a1a35	GH-103727: Avoid advancing tokenizer too far in f-string mode (GH-103775)	2023-04-24 12:30:21 -06:00
Lysandros Nikolaou	05b3ce7339	GH-103718: Correctly cache and restore f-string buffers when needed (GH-103719)	2023-04-23 13:06:10 -06:00
Pablo Galindo Salgado	d4aa8578b1	gh-102856: Clean some of the PEP 701 tokenizer implementation (#103634 )	2023-04-19 14:51:31 -06:00
Pablo Galindo Salgado	1ef61cf71a	gh-102856: Initial implementation of PEP 701 (#102855 ) Co-authored-by: Lysandros Nikolaou <lisandrosnik@gmail.com> Co-authored-by: Batuhan Taskaya <isidentical@gmail.com> Co-authored-by: Marta Gómez Macías <mgmacias@google.com> Co-authored-by: sunmy2019 <59365878+sunmy2019@users.noreply.github.com>	2023-04-19 11:18:16 -05:00
Pablo Galindo Salgado	417206a05c	gh-99891: Fix infinite recursion in the tokenizer when showing warnings (GH-99893) Automerge-Triggered-By: GH:pablogsal	2022-11-30 03:36:06 -08:00
Pablo Galindo Salgado	e13d1d9dda	gh-99581: Fix a buffer overflow in the tokenizer when copying lines that fill the available buffer (#99605 )	2022-11-20 20:20:03 +00:00
Victor Stinner	4ce2a202c7	gh-99300: Use Py_NewRef() in Parser/ directory (#99330 ) Replace Py_INCREF() with Py_NewRef() in C files of the Parser/ directory and in the PEG generator.	2022-11-10 15:30:05 +01:00
Lysandros Nikolaou	3de08ce8c1	gh-97997: Add col_offset field to tokenizer and use that for AST nodes (#98000 )	2022-10-07 14:38:35 -07:00
Lysandros Nikolaou	cbf0afd8a1	gh-97973: Return all necessary information from the tokenizer (GH-97984) Right now, the tokenizer only returns type and two pointers to the start and end of the token. This PR modifies the tokenizer to return the type and set all of the necessary information, so that the parser does not have to this.	2022-10-06 16:07:17 -07:00
Pablo Galindo Salgado	aab01e3524	gh-96670: Raise SyntaxError when parsing NULL bytes (#97594 )	2022-09-27 23:23:42 +01:00
Matthias Görgens	81e36f350b	gh-96678: Fix UB of null pointer arithmetic (GH-96782) Automerge-Triggered-By: GH:pablogsal	2022-09-13 06:14:35 -07:00
Michael Droettboom	8bc356a7dd	gh-96268: Fix loading invalid UTF-8 (#96270 ) This makes tokenizer.c:valid_utf8 match stringlib/codecs.h:decode_utf8. It also fixes an off-by-one error introduced in 3.10 for the line number when the tokenizer reports bad UTF8.	2022-09-07 14:23:54 -07:00
Michael Droettboom	05692c67c5	gh-96611: Fix error message for invalid UTF-8 in mid-multiline string (#96623 )	2022-09-07 00:12:16 +01:00
Pablo Galindo Salgado	36fcde61ba	gh-94360: Fix a tokenizer crash when reading encoded files with syntax errors from stdin (#94386 ) * gh-94360: Fix a tokenizer crash when reading encoded files with syntax errors from stdin Signed-off-by: Pablo Galindo <pablogsal@gmail.com> * nitty nit Co-authored-by: Łukasz Langa <lukasz@langa.pl>	2022-07-05 17:39:21 +01:00
Serhiy Storchaka	6fd4c8ec77	gh-93741: Add private C API _PyImport_GetModuleAttrString() (GH-93742) It combines PyImport_ImportModule() and PyObject_GetAttrString() and saves 4-6 lines of code on every use. Add also _PyImport_GetModuleAttr() which takes Python strings as arguments.	2022-06-14 07:15:26 +03:00
Kumar Aditya	cb04a09d2d	GH-93207: Remove HAVE_STDARG_PROTOTYPES configure check for stdarg.h (#93215 )	2022-05-27 13:30:45 +02:00
Victor Stinner	5115a16831	gh-93103: Parser uses PyConfig.parser_debug instead of Py_DebugFlag (#93106 ) * Replace deprecated Py_DebugFlag with PyConfig.parser_debug in the parser. * Add Parser.debug member. * Add tok_state.debug member. * Py_FrozenMain(): Replace Py_VerboseFlag with PyConfig.verbose.	2022-05-24 22:35:08 +02:00
Victor Stinner	da5727a120	gh-92651: Remove the Include/token.h header file (#92652 ) Remove the token.h header file. There was never any public tokenizer C API. The token.h header file was only designed to be used by Python internals. Move Include/token.h to Include/internal/pycore_token.h. Including this header file now requires that the Py_BUILD_CORE macro is defined. It no longer checks for the Py_LIMITED_API macro. Rename functions: * PyToken_OneChar() => _PyToken_OneChar() * PyToken_TwoChars() => _PyToken_TwoChars() * PyToken_ThreeChars() => _PyToken_ThreeChars()	2022-05-11 23:22:50 +02:00
Serhiy Storchaka	43a8bf1ea4	gh-87999: Change warning type for numeric literal followed by keyword (GH-91980) The warning emitted by the Python parser for a numeric literal immediately followed by keyword has been changed from deprecation warning to syntax warning.	2022-04-27 20:15:14 +03:00
Christian Heimes	3df0e63aab	bpo-46315: Use fopencookie only on Emscripten 3.x and newer (GH-32266)	2022-04-02 23:11:38 +02:00
Hugo van Kemenade	6881ea936e	bpo-47126: Update to canonical PEP URLs specified by PEP 676 (GH-32124)	2022-03-30 12:00:27 +01:00
Christian Heimes	9b889b5bda	bpo-46315: Use fopencookie() to avoid dup() in _PyTokenizer_FindEncodingFilename (GH-32033) WASI does not have dup() and Emscripten's emulation is slow.	2022-03-22 17:08:51 +01:00
Oleg Iarygin	13b0412223	bpo-46920: Remove code that has explainers why it was disabled (GH-31813)	2022-03-14 17:04:22 +01:00
Serhiy Storchaka	090e5c4b94	bpo-46820: Fix a SyntaxError in a numeric literal followed by "not in" (GH-31479) Fix parsing a numeric literal immediately (without spaces) followed by "not in" keywords, like in "1not in x". Now the parser only emits a warning, not a syntax error.	2022-02-22 09:51:51 +02:00

1 2 3 4 5 ...

346 Commits