Repository navigation
Add "strict" mode to ignore () and [] when building transcript #441
Description
Activity
This is intentional, meant to address the fact that many caption/subtitle files use parentheses and square brackets to identify speaker names. When it encounters these characters, that forces a line break and displays the content inside those characters in a span with class="able-unspoken".
I admit this isn't ideal. One solution would be to add support for data-vtt-mode="strict", in which case all parsing would be based solely on WebVTT rules.
Hi Terrill,
Support for data-vtt-mode="strict" would be great. Does that include support for Bold, Underline, Initialize, and color tags? I know the clients in the writing and linguistics department would be happy.
Able Player already supports bold, italics, and underline, both in the captions and transcript. Any other styling will be addressed when we get to Issue #13
- changed the title
[-]Multiple cues on one line causing new lines[/-][+]Add "strict" mode to ignore () and [] when building transcript[/+]on Dec 21, 2020 Is there an escape character that would allow displaying [ ] without the line break and bolding?
Reacted by Vinícius K-Max and Deborah PickettRather than create a new issue I just want to mention that
data-lyrics-modeadds unwanted line breaks around all styling tags like<i>and<b>. This seems similar to the OP on this one.
Probably this was missed because the examples do not usedata-lyrics-mode, e.g. this one where the bolding is not causing such problems.Might be worth an issue depending on how things shake out. The project is pretty stagnant right now. Would be good if a bunch of us users could chip in more. I'm including myself in that!
The bug I just mentioned seemed different enough that I decided to make a new issue and dig into it.
We have an author/creator on fulcrum.org who noticed this problem in his video files straight away, given we use
data-lyrics-modeeverywhere with Able Player.aside: I didn't see any support for underline in the transcript code, though it would be easy to add if that should be done.
I'm not going to add support for underline, I think. The underline is a very strong signal for links, and having it as a formatting-only tag is undesirable.
Adding this to the 4.9 milestone. I think that the special handling for parentheses and brackets needs some better handling - strict mode may be sufficient, but I think it's a problem right now that any parenthetical in spoken text triggers this extra markup.
- added a commit that references this issue
on Jun 27, 2026 Fixed by 751ae54
At this point, defaults to false; but I'm going to raise a separate issue to switch it to true in the future.
We have noticed that if a transcript has multiple cues on one line a new line will be generated starting with the second cue. For example:
00:05:10.000 --> 00:05:16.000
Sí. Entonces la (OL) playlist (OL) era de- ¿y tú cuándo empezaste a escuchar a [angry] Julieta, y a Natalia, y a Elefante a-?
will display on the transcript box as:
(INV) Sí. Entonces la (OL) playlist
(OL) era de- ¿y tú cuándo empezaste a escuchar a
[angry] Julieta, y a Natalia, y a Elefante a-?
instead of:
(INV) Sí. Entonces la (OL) playlist (OL) era de- ¿y tú cuándo empezaste a escuchar a [angry] Julieta, y a Natalia, y a Elefante a-?
it doesn't matters what type of cue it is including HTML tags for bold, underline, initialize.